arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

解读环境:大语言模型中规范能力的基础、设计与挑战

Reading the Room: Foundations, Design, and Challenges of Normative Competence in LLMs

Andrea Wynn, Harsh Satija, Seokhyun, Baek, Anqi Liu, Eric Nalisnick, Gillian K. Hadfield

arXiv 2610.10906首次发表:更新:

发表机构

Johns Hopkins University; Vector Institute(约翰斯·霍普金斯大学; 矢量研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究引入多智能体辩论设置研究大语言模型的规范能力,发现基线模型难学规范、泛化性不足且存在无选择性归因失败,证明当前AI擅长行为模仿但缺乏辨别社会秩序的能力。

AI 中文摘要

人类社会由规范体系治理:即产生“规范”的共享标准,这些规范规定了可接受的行为,并通过社区制裁实施。让日益自主的AI系统与这些规范对齐是核心的对齐挑战,而规范数量庞大、变化迅速且往往具有任意性(如着装或语言惯例)的事实让这一挑战更为复杂。因此,对齐需要“规范能力”:仅通过交互就能辨别社区实施的规范,而无需依赖静态预训练知识的能力。我们引入了一种多智能体社区辩论设置,其中辩论的访问受合成规范管控,以在脱离预训练暴露的情况下单独研究规范能力。我们发现,即使遵循规范能提升准确率,基线大语言模型智能体也无法学习规范。随后,我们对各种“规范模块”——用于规范推理的架构组件——进行实验,发现遵循规范的表现对规范的类型和支撑规范模块的模型能力高度敏感,这表明存在泛化性不足的问题。此外,当真实规范伴随特殊的非规范行为时,大语言模型智能体会表现出无选择性归因失败:它们会不加区分地复制特殊噪声与实施的规则,即使模仿不必要的行为会被明确惩罚,这种模式依然存在。据我们所知,本研究是首个将大语言模型中的规范能力付诸实施并进行评估的工作,证明当前AI系统擅长行为模仿,但缺乏辨别社会实施秩序的能力。

英文摘要

Human communities are governed by normative systems: shared standards that produce \textit{norms} dictating acceptable behavior, enforced through community sanctioning. Aligning increasingly autonomous AI systems with these norms is a central alignment challenge, complicated by the fact that norms are vast in number, change quickly, and are often arbitrary (e.g., dress or language conventions). Thus, alignment requires \textit{normative competence}: the ability to discern from interaction alone what norms a community enforces without relying on static pretrained knowledge. We introduce a multi-agent community debate setting, where access to debate is governed by synthetic norms, to study normative competence in isolation from pretraining exposure. We show that baseline LLM agents fail to learn norms even when doing so would improve their accuracy. We then experiment with various \textit{normative modules} -- architectural components for norm inference -- finding that norm-following is highly sensitive to both the style of norm and the model powering the normative module, suggesting a lack of generalizability. Furthermore, when idiosyncratic, non-normative behaviors accompany the true norm, LLM agents exhibit an unselective attribution failure: they indiscriminately copy idiosyncratic noise alongside enforced rules, a pattern that persists even when imitating unnecessary behaviors is explicitly penalized. To the best of our knowledge, our work is the first to operationalize and evaluate normative competence in LLMs, demonstrating that current AI systems excel at behavioral mimicry but lack the capacity to discern socially enforced order.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑