arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

评估与改进大语言模型(LLM)的自我建模能力

Evaluating and Improving LLM Self-Modeling

Siqi Zeng, Andre N. Assis, Rowan Wang

arXiv 2608.30980首次发表:更新:

发表机构

Constellation; Anthropic(星群公司; Anthropic公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对LLM的自我建模能力,构建基准测试评估其水平,开发合成数据管道结合强化学习提升该能力,发现提升未体现一致内省。

AI 中文摘要

我们研究自我建模:大语言模型(LLM)回答关于自身行为问题的能力,重点关注可验证的行为问题,比如提示词编辑是否会改变模型的最终答案。为衡量该能力,我们推出一项基准测试,涵盖各类自我建模问题。当前模型展现出可观但有限的自我建模技能,且在关于自身行为的简单反事实问题上存在系统性错误。为提升自我建模技能,我们开发可扩展的合成数据管道,生成自我建模训练数据,并证明强化学习可提升三大开源模型系列的整体自我建模技能,且部分可迁移至保留任务。不过,这些提升似乎并未构成一致的内省:自我建模的改进可能并非源于对模型内部决策过程的特权访问。

英文摘要

We study self-modeling: an LLM's ability to answer questions about its own behavior. We focus on verifiable behavioral questions, such as whether a prompt edit would change the model's final answer. To measure this capability, we introduce a benchmark that tests diverse types of self-modeling questions. Current models show non-trivial but limited self-modeling skill, and make systematic mistakes on simple counterfactual questions about their own behavior. To improve self-modeling skill, we develop a scalable synthetic-data pipeline that produces self-modeling training data, and show that reinforcement-learning can improve aggregate self-modeling skill across three open-source model families with some transfer to held-out tasks. These gains, however, do not seem to constitute introspection consistently: improved self-modeling may not arise from privileged access to the model's internal decision process.

Comments89 pages, 25 figures. Published as a conference paper at EMNLP '26 (Main)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑