发表机构
Predictably Weird(Predictably Weird)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过九项行为评估测试语言模型的自我认知,发现其自我报告预测能力弱且不具自我特异性,第一人称框架带来奉承偏差,微调收益有限,表明模型自我表述实为泛化理论加偏见。
AI 中文摘要
语言模型能够流畅地描述它们会如何表现:是否会屈服于反驳、是否会滥用工具、是否会在压力下撒谎。这种描述真的关乎正在说话的模型本身吗?我们将自我认知转化为一项预测测试。在九项行为评估中,我们测量模型在不同条件下的行为表现,要求其预测这些比率,并将其预测与去除自我因素的对照问题进行比较。我们发现:(i)直接自我报告很弱(r = +0.04),即使向模型展示确切的项目,也仅将预测提升至+0.24。关键在于,关于“一般的有能力的AI智能体”的相同项目知情问题表现同样好(+0.28),而其他模型关于自身的回答对目标模型的预测效果至少与其自身回答相当。(ii)前沿模型规模并未显著改变这一模式:预测上的任何提升都不是自我特异的,并且与更好的AI助手行为理论一致,而非更好的自我认知。(iii)第一人称框架确实有一个稳健的影响:它使报告偏向奉承方向,相对于关于一般智能体的相同问题,低估了有害行为。(iv)在模型自身行为记录上进行微调可以教会狭窄的自我预测,但它也改变了被预测的行为,且提升不能广泛迁移。实际含义很简单:询问模型它会做什么,主要揭示的是关于一般AI助手的理论,加上一种有利的偏见,而非该模型的专属知识。
英文摘要
Language models can fluently describe how they would behave: whether they would cave to pushback, misuse a tool, or lie under pressure. Is that description actually about the model speaking? We turn self-knowledge into a prediction test. Across nine behavioral evaluations, we measure how a model behaves under different conditions, ask it to predict those rates, and compare its predictions with controls that remove the self from the question. We find that: (i) Direct self-report is weak (r = +0.04), and even showing the model the exact items only raises prediction to +0.24. Crucially, the same item-informed question about "capable AI agents in general" does just as well (+0.28), while other models' answers about themselves predict the target model at least as well as its own. (ii) Frontier scale does not detectably change this pattern: any gains in prediction are not self-specific, and are consistent with a better theory of how AI assistants behave rather than better self-knowledge. (iii) First-person framing does have one robust effect: it shifts reports in the flattering direction, understating harmful behavior relative to the same question about a generic agent. (iv) Finetuning on a model's own behavioral record can teach narrow self-predictions, but it also changes the behavior being predicted and the gains do not transfer broadly. The practical implication is simple: asking a model what it would do mostly reveals a theory of AI assistants in general, plus a favorable bias, rather than privileged knowledge of that model.