arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

开放权重掩码内省:测量语言模型可报告自身计算的能力

Open-Weight Masked Introspection: Measuring What Language Models Can Report About Their Own Computation

Emilio Ferrara

arXiv 2608.20569首次发表:更新:

发表机构

University of Southern California(南加州大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究构建OWMI框架检验8个开放权重模型的内省能力,发现其无法区分真实干预与虚假运行,仅内部存在相关信息,失败源于内部状态到文本报告的路径,需对照内部参考验证模型证词。

AI 中文摘要

前沿模型是否具备对自身内部状态进行内省的能力?近期研究表明,在特定条件下,足够复杂的模型可对自身内部进行审计、指出变化并自信地报告结果。我们针对来自7个家族的8个开放权重模型检验了该主张,结果发现这些模型不具备此类能力:当被问及自身计算是否被改变时,所有模型的回答准确率均未超过随机水平。为开展检验,我们构建了Open-Weight Masked Introspection(OWMI,开放权重掩码内省)框架,该框架会对残差流位点、注意力头及稀疏自编码器特征进行干预,随后对照三类基准条件询问模型关于该变化的情况:无任何改变的虚假运行、与影响匹配的随机扰动,以及仅可见输出的纯文本观测器。在超过78000次测量中,没有任何模型的报告能以超过随机水平的能力区分真实干预与虚假运行(AUROC约为0.5007),且等价检验将效应值限定在AUROC的0.15个百分点以下。令人惊讶的是,所有所需信息均存在于模型内部:针对此类干预进行微调的模型能在保留方向上实现近乎完美的恢复,线性探针从相同激活中恢复干预存在的准确率为75%至95.8%,在模型输出前的最后一层,准确率提升至无保留误差。其中一个模型的信号体现在置信度而非文本上:其是/否报告从未变化,但附带的置信度以AUROC 0.647的水平区分了干预与虚假运行。这种失败存在于从内部状态到文本报告的路径中,因此,读取模型自身证词的监督需要对照内部参考进行验证。尽管我们的结果表明当前开放权重模型不具备内省能力,但对于未来模型,相关争论仍未定论。

英文摘要

Are frontier models able to introspect about their internal states? Recent work suggests that under certain conditions a complex enough model can audit its own internals, call out what changed, and report back confidently about it. We tested that claim on eight open-weight models from seven families and found no such ability: asked whether their own computation had been altered, none answered better than chance. To test it we built Open-Weight Masked Introspection (OWMI), a framework that intervenes on residual-stream sites, attention heads and sparse-autoencoder features, then interrogates the model about the change against the null conditions an answer has to beat: sham runs where nothing was altered, impact-matched random perturbations, and a text-only observer that sees only the visible output. Over 78,000 measurements, no model's report discriminates a real intervention from a sham beyond chance (AUROC ~0.5007), and an equivalence test bounds the effect below 0.15 percentage points of AUROC. Surprisingly, all the information needed is in the models. A model fine-tuned to report this class of intervention reaches near-perfect recovery on held-out directions, and a linear probe recovers intervention presence from the same activations at 75% to 95.8% accuracy, sharpening to no held-out error at the last layer before the model speaks. In one model the signal surfaces in the confidence rather than the words: its yes-or-no report never varies, while the confidence attached to it separates intervention from sham at AUROC 0.647. The failure sits in the path from internal state to verbal report, so oversight that reads a model's own testimony needs validating against an internal reference. While our results show the inability of current open-weight models to introspect, the debate is not settled for future models.

CommentsWe release OWMI as a library so that this emerging ability can be measured as it develops. Hugging Face OWMI library: https://huggingface.co/emilioferrara/owmi

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑