发表机构
Google Research; Cornell University; Stanford University; Google DeepMind(谷歌研究院; 康奈尔大学; 斯坦福大学; 谷歌DeepMind)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对语言模型融入生活引发的长期人机交互风险,结合社会科学测量与NLP计算方法,提出通过长期测量建模人类行为变化,实现问题行为在线检测以缓解用户长期风险。
AI 中文摘要
语言模型凭借其“类人性”及快速融入用户日常生活的特点,已成为一种全新类型的技术。这种特性组合会引发纵向风险——即人类的认知、发展及社会情感变化,这类变化可能不会在短期交互中显现,却会对用户产生持久的长期影响。这为自然语言处理(NLP)领域确立了一项关键的新任务:从对文本生成的静态、短期评估,转向对行为变化的长期测量,以实现对人机交互的历时性理解。在本研究中,我们借鉴了社会科学领域用于理解纵向数据中涌现现象的核心测量方法,探讨了NLP领域的计算方法需如何与这些测量相结合,不仅用于理解人机交互的长期安全风险,还能助力引导模型开发,使其为用户带来积极而非消极的结果。这种将人类行为变化建模为模型交互函数的能力,可实现对问题行为的在线而非事后检测,应被应用于对齐框架中,以缓解用户面临的长期风险。
英文摘要
Language models have taken on the role of a very new type of technology, by virtue of their "human-ness" and rapid integration into users' daily lives. This combination of features can introduce longitudinal risks---cognitive, developmental and socio-affective changes in humans---that might not surface during a short-term interaction, but can have lasting long-term effects on users. This forms the basis of a critical new mission for NLP: to pivot from static, short-term evaluations of text generations to long-term measurements of behavioral changes, towards a diachronic understanding of human-model interactions. In this work, we draw from measurements used in social science fields that are crucial to understand emergent phenomena in longitudinal data. We discuss how computational methods in the field of NLP need to be combined with such measurements, not only to understand long-term safety risks of human-model interactions, but to help steer model development towards positive rather than negative outcomes for users. This ability to model human behavioral shifts as a function of model interactions can facilitate online rather than post-hoc detection of problematic behaviors, and should be leveraged in alignment frameworks to mitigate long-term risks in users.