发表机构
Constellation; Anthropic(星群公司; Anthropic公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对LLM的自我建模能力,构建基准测试评估其水平,开发合成数据管道结合强化学习提升该能力,发现提升未体现一致内省。
AI 中文摘要
我们研究自我建模:大语言模型(LLM)回答关于自身行为问题的能力,重点关注可验证的行为问题,比如提示词编辑是否会改变模型的最终答案。为衡量该能力,我们推出一项基准测试,涵盖各类自我建模问题。当前模型展现出可观但有限的自我建模技能,且在关于自身行为的简单反事实问题上存在系统性错误。为提升自我建模技能,我们开发可扩展的合成数据管道,生成自我建模训练数据,并证明强化学习可提升三大开源模型系列的整体自我建模技能,且部分可迁移至保留任务。不过,这些提升似乎并未构成一致的内省:自我建模的改进可能并非源于对模型内部决策过程的特权访问。
英文摘要
We study self-modeling: an LLM's ability to answer questions about its own behavior. We focus on verifiable behavioral questions, such as whether a prompt edit would change the model's final answer. To measure this capability, we introduce a benchmark that tests diverse types of self-modeling questions. Current models show non-trivial but limited self-modeling skill, and make systematic mistakes on simple counterfactual questions about their own behavior. To improve self-modeling skill, we develop a scalable synthetic-data pipeline that produces self-modeling training data, and show that reinforcement-learning can improve aggregate self-modeling skill across three open-source model families with some transfer to held-out tasks. These gains, however, do not seem to constitute introspection consistently: improved self-modeling may not arise from privileged access to the model's internal decision process.
Comments89 pages, 25 figures. Published as a conference paper at EMNLP '26 (Main)