arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2602.10625cs.AIcs.CL

思考还是不思考,这是大型推理模型在心智理论任务中的问题

To Think or Not To Think, That is The Question for Large Reasoning Models in Theory of Mind Tasks

  • Arizona State University, Tempe, Arizona, United States(亚利桑那州立大学)
  • Microsoft Research Asia, Beijing, China(微软亚洲研究院)

机构由 AI 辅助整理,请以论文原文为准。

Nanxu Gong, Haotian Li, Sixun Dong, Jianxun Lian, Yanjie Fu, Xing Xie

更新

AI总结:

本文研究了大型推理模型在心智理论任务中的表现,发现其并不总优于非推理模型,且依赖选项匹配而非真实推理,提出干预方法以缓解问题。

AI中文摘要:

心智理论(ToM)评估模型是否能够推断出隐藏的心理状态,如信念、欲望和意图,这对于自然的社会互动至关重要。尽管最近在大型推理模型(LRMs)上的进步提高了数学和编程中的逐步推理能力,但尚不清楚这种优势是否能转移到社会认知技能上。我们对九种先进的大型语言模型(LLMs)进行了系统研究,比较推理模型与非推理模型在三个代表性的ToM基准上的表现。结果表明,推理模型并不总是优于非推理模型,有时甚至表现更差。细致分析揭示了三个见解。首先,慢思考崩溃:随着响应变长,准确性显著下降,且较大的推理预算损害性能。第二,适度和适应性的推理有助于表现:限制推理长度可以缓解失败,而不同的成功模式表明动态适应的必要性。第三,选项匹配捷径:当多个选择选项被移除时,推理模型显著提高,表明依赖于选项匹配而非真正的推理。我们还设计了两种干预方法:Slow-to-Fast(S2F)适应性推理和Think-to-Match(T2M)捷径预防,以进一步验证和缓解这些问题。所有结果表明,LRMs在正式推理(如数学、代码)上的进步不能完全转移到ToM,这是社会推理中的典型任务。我们得出结论,实现稳健的ToM需要开发超出现有推理方法的独特能力。

英文摘要:

Theory of Mind (ToM) assesses whether models can infer hidden mental states such as beliefs, desires, and intentions, which is essential for natural social interaction. Although recent progress in Large Reasoning Models (LRMs) has boosted step-by-step inference in mathematics and coding, it is still underexplored whether this benefit transfers to socio-cognitive skills. We present a systematic study of nine advanced Large Language Models (LLMs), comparing reasoning models with non-reasoning models on three representative ToM benchmarks. The results show that reasoning models do not consistently outperform non-reasoning models and sometimes perform worse. A fine-grained analysis reveals three insights. First, slow thinking collapses: accuracy significantly drops as responses grow longer, and larger reasoning budgets hurt performance. Second, moderate and adaptive reasoning benefits performance: constraining reasoning length mitigates failure, while distinct success patterns demonstrate the necessity of dynamic adaptation. Third, option matching shortcut: when multiple choice options are removed, reasoning models improve markedly, indicating reliance on option matching rather than genuine deduction. We also design two intervention approaches: Slow-to-Fast (S2F) adaptive reasoning and Think-to-Match (T2M) shortcut prevention to further verify and mitigate the problems. With all results, our study highlights the advancement of LRMs in formal reasoning (e.g., math, code) cannot be fully transferred to ToM, a typical task in social reasoning. We conclude that achieving robust ToM requires developing unique capabilities beyond existing reasoning methods.

↑