发表机构
StackOne Technologies(StackOne 科技公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出内部识别探针(PIR)方法,借鉴法医隐蔽信息测试,通过读取模型内部状态识别其隐藏的已知答案,以区分模型不愿回答与不能回答,支持表现不佳审计和遗忘验证。
AI 中文摘要
大型语言模型可以持有它们不会报告的知识。一个模型可能在能力评估中表现不佳(sandbagging),或违背其内部所知来作答,而仅凭其输出无法判断它是在隐藏答案还是根本没有答案。我们借鉴了隐蔽信息测试(Concealed Information Test),这是一种法医学方法,通过向嫌疑人呈现真实细节与似是而非的干扰项,并测量其对所识别项目的更强反应来识别犯罪知识。我们的方法——内部识别探针(Probe of Internal Recognition, PIR),在模型内部执行同样的操作。它提出一个问题及其候选答案,并从模型的内部状态中读取模型识别为正确的候选答案。PIR无需参考,不需要诚实的参考模型,也不需要标注的真实语料库。在来自五个模型家族(Gemma、Qwen、Llama、Mistral和Phi)的八个模型上,PIR以0.70至0.87的平衡准确率恢复被识别的答案,远高于0.28至0.40的未知项目基线和0.25的随机概率。在我们测试的各种隐藏形式下,从提示欺骗和训练性表现不佳到外部密码锁定和电路断开的检查点,它都保持可读性,识别率在0.85至0.93之间。当模型隐藏已知答案时,识别率保持较高。当通过遗忘(unlearning)移除知识时,识别率下降到模型从未知道的问题的水平。因此,PIR区分了一个不愿回答的模型和一个不能回答的模型,这支持了表现不佳审计和遗忘验证。该信号是因果性的,在黑盒行为线索之外增加了信息,并从多项选择题扩展到自由形式的生成。
英文摘要
Large language models can hold knowledge they do not report. A model may sandbag on a capability evaluation, or answer against what it internally knows, and its outputs alone cannot tell whether it is hiding an answer or simply does not have one. We borrow the Concealed Information Test, a forensic method that identifies guilty knowledge by presenting a suspect with the true detail among plausible decoys and measuring a stronger response to the item they recognize. Our method, Probe of Internal Recognition (PIR), does the same inside a model. It presents a question with its candidate answers and reads, from the model's internal states, which candidate the model recognizes as correct. PIR is reference-free, needing no honest reference model and no labeled truth corpus. Across eight models from five families (Gemma, Qwen, Llama, Mistral, and Phi), PIR recovers the recognized answer at 0.70 to 0.87 balanced accuracy, well above the 0.28 to 0.40 unknown-item baseline and the 0.25 chance rate. It stays readable across every form of concealment we test, from prompted deception and trained sandbagging to external password-locked and circuit-broken checkpoints, with recognition between 0.85 and 0.93. When the model hides a known answer, recognition stays high. When unlearning removes the knowledge, recognition drops to the level of a question the model never knew. PIR therefore separates a model that will not answer from one that cannot, which supports sandbagging audits and unlearning verification. The signal is causal, adds information beyond black-box behavioral cues, and extends from multiple-choice questions to free-form generation.