发表机构
Leiden University; LIACS(莱顿大学; 莱顿高级计算机科学研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究考察经强化学习训练的推理LLMs在ToM任务中的表现,发现其对提示与任务扰动的鲁棒性提升,支持鲁棒性解释而非ToM特有的新能力。
AI 中文摘要
大型语言模型(LLMs)近期在心理理论(Theory of Mind, ToM)测试中展现出强劲性能,引发了关于其底层能力本质与有效性的争论。与此同时,通过带可验证奖励的强化学习训练的面向推理的LLMs,在一系列基准测试中取得了显著提升。本研究采用机器心理实验的新适配方法及现有基准结果,考察此类推理模型在ToM任务中的行为,发现推理模型对提示变化和任务扰动的鲁棒性持续增强。分析表明,这些提升至少部分源于模型在提示与任务变化下更易得到正确答案,该结果支持基于鲁棒性的解释,而非ToM特有的新能力。
英文摘要
Large language models (LLMs) have recently shown strong performance on Theory of Mind (ToM) tests, prompting debate about the nature and validity of the underlying capabilities. At the same time, reasoning-oriented LLMs trained via reinforcement learning with verifiable rewards have demonstrated notable improvements across a range of benchmarks. In this work, we examine the behavior of such reasoning models in ToM tasks using novel adaptations of machine psychological experiments together with results from established benchmarks. We observe that reasoning models consistently exhibit increased robustness to prompt variations and task perturbations. Our analysis suggests these gains come at least partly from models being more robust at reaching the correct answer under prompt and task variation. We read this as evidence for a robustness-based account rather than for a new ToM-specific ability.
CommentsAccepted for 29th International Conference on Discovery Science, October 5-9, 2026, Mainz, Germany