arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.12022cs.CLcs.AIcs.LG

探究大语言模型推理中的社会归因:一种理论引导的探测方法

Examining Social Attribution in LLM Reasoning: A Theory-Guided Probing Methodology

Zhaoxin Yu, Qingchao Kong, Dajun Zeng, Wenji Mao

首次发表
浏览论文内容

中文总结 AI 辅助

本文首次系统探究LLM的社会归因,构建含7639个问题的基准,评估32个LLM和5个基线,发现LLM与人类判断部分一致且一致性随模型规模提升,部分归因维度可从隐藏状态解码。

中文摘要 AI 辅助

大语言模型(LLMs)正越来越多地部署在社会技术系统中,其中社会归因(即将外部事件归因于主体社会行为的原因和理由的推理过程)发挥着关键作用。这些过程涉及对社会原因、责任以及对主体的责备/赞誉的判断。尽管归因模型在社会心理学和认知领域通过归因理论得到了充分研究,但社会归因在人工智能领域,尤其是LLM的社会推理中仍未得到充分探索。本文首次对LLM的社会归因进行了系统探索,研究聚焦于责任和责备归因,考察当前LLM的判断及其潜在内部机制。我们在归因理论的指导下构建了一个社会归因基准,该基准包含基于归因理论研究经典场景的小插曲(Vignette)子集,以及基于现实世界社会叙事的现实(Reality)子集,共生成7639个责任/责备判断问题。在此基础上,我们评估了32个代表性LLM和5个基础非LLM基线。为进一步探究LLM判断过程的潜在内部机制,我们开发了一种基于探测的方法,用于研究5个关键归因维度的隐空间表示,以及它们对LLM判断的影响与人类社会归因影响的一致性。我们的研究结果表明,当前LLM与人类的责任和责备判断存在可测量但不完全的一致性,且这种一致性与模型规模呈正相关。部分归因维度可从LLM隐藏状态的特定位置系统地解码,且它们对最终判断的影响与人类归因理论所指示的影响一致。该数据集及相关代码可在指定URL获取。

英文摘要

Large language models (LLMs) are increasingly deployed in sociotechnical systems where social attribution, the reasoning process attributing external events to the causes and reasons of agents' social behaviors, plays a critical role. These processes involve judgments of social cause, responsibility, and blame/credit to agents. Although attributional models are well-studied in social psychology and cognition through Attribution Theory, social attribution remains underexplored in AI, particularly LLM social reasoning. This paper provides the first systematic exploration of LLM social attribution. Our work focuses on responsibility and blame attributions, examining current LLMs' judgments and their underlying internal mechanisms. Guided by attribution theory, we construct a social attribution benchmark consisting of a Vignette subset based on classic scenarios from attribution theory research and a Reality subset based on real-world social narratives, yielding 7,639 responsibility/blame judgment questions. On this basis, we evaluate 32 representative LLMs and 5 basic non-LLM baselines. To further explore the internal mechanisms underlying the LLM judgment process, we develop a probing-based methodology to investigate the latent-space representations of 5 key attribution dimensions and the consistency of their influences on LLM judgments compared to those in human social attribution. Our research findings reveal that current LLMs exhibit measurable but incomplete agreement with human responsibility and blame judgments, and meanwhile, this agreement is positively correlated with model size. Some attribution dimensions are systematically decodable from specific positions in LLM hidden states, and their influences on the final judgment are consistent with those indicated by human Attribution Theory. The dataset and associated code are available at https://github.com/Yuzhaoxin946/SAB-Bench.

发表机构

  • Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
  • University of Chinese Academy of Sciences(中国科学院大学)

机构由 AI 辅助整理,请以论文原文为准。

↑