arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.05437cs.AIcs.CY

超越对与错:评估大型语言模型中的二阶社会推理

Beyond Right and Wrong: Evaluating Second-order Social Reasoning in Large Language Models

  • University of Pennsylvania(宾夕法尼亚大学)
  • The World Bank(世界银行)

机构由 AI 辅助整理,请以论文原文为准。

Sunny Rai, Jinyi Kuang, Reyhan Jamalova, Annie Lou, Cristina Bicchieri, Niyati Malhotra, Victor Hugo Orozco-Olvera, Ana Maria Munoz-Boudet, Lyle H Ungar, Sharath C Guntuku

AI总结:

本研究提出评估LLM元规范推理的框架与NormReact数据集,发现模型过度预测惩罚,低估人类宽容与关系校准,导致社会调节图景失真。

AI中文摘要:

以往的AI对齐工作主要聚焦于一阶社会规范——教导模型什么是社会可接受或不可接受的(例如,“不要偷窃”)。然而,社会智能不仅依赖于规范识别,还依赖于预测谁会执行规范以及如何执行(例如,公开羞辱甚至监禁)。这些二阶期望,即元规范,支配着当社会规则被打破时人们如何回应。我们引入了一个新颖的框架,用于评估大型语言模型(LLMs)中的元规范推理,该框架沿两个维度展开:情绪评估和行为反应,并提出了新的分类任务,即预测违规者的自我调节和观察者的他人调节。我们发布了一个多视角数据集NormReact,包含450个规范违规场景,针对违规者性别和观察者社会亲近度,对情绪和行为反应进行了人工标注。当前的LLMs描绘了一个更严酷的社会世界:在六个模型中,它们过度预测了人类预期不作为情况下的负面制裁,并且随着社会距离的增加,与人类判断的一致性会恶化。这些发现表明,在从冲突调解到政策模拟等规范敏感领域中的AI系统,可能面临产生扭曲的社会调节图景的风险:一种过度代表惩罚而低估了宽容、克制和关系校准的图景,而这些正是现实世界中规范执行的特征。

英文摘要:

Previous AI alignment efforts have focused primarily on first-order social norms -- teaching models what is socially acceptable or unacceptable (e.g., `do not steal'). However, social intelligence depends not only on norm recognition, but also on anticipating who will enforce it and how (e.g., public shame or even imprisonment). These second-order expectations, known as metanorms, govern how people respond when social rules are broken. We introduce a novel framework for evaluating metanorm reasoning in Large Language Models (LLMs) along two dimensions: emotional appraisal and behavioral response, and propose new classification tasks, namely, predicting self-regulation in violators, and other-regulation in observers. We release a multi-perspective dataset, NormReact, of 450 norm violation scenarios, hand-annotated for emotions and behavioral responses across norm violators' gender and observers' social closeness. Current LLMs portray a harsher social world: across six models, they overpredict negative sanctions where humans would expect inaction, and alignment with human judgments deteriorates as social distance increases. These findings suggest that AI systems in norm-sensitive domains from conflict mediation to policy simulation, may risk producing a distorted picture of social regulation: one that over-represents punishment and under-represents the tolerance, restraint, and relational calibration that characterize actual norm enforcement in real world.

↑