arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.10739cs.LGcs.AIcs.CL

真相从未消失:顺从上下文真值探针中的完全混叠

The Truth Was Never Gone: Perfect Aliasing in Compliant-Context Truth Probes

Dylan Jayabahu

首次发表
浏览论文内容

中文总结 AI 辅助

本研究揭示顺从上下文真值探针存在完全混叠问题,通过混合拟合分离真值与规定动作,在对抗上下文中将AUROC从近零提升至1.000,证明线性可恢复性。

中文摘要 AI 辅助

在真实报告与任务规定动作相重合的位置拟合的真值探针,无法仅凭其拟合标签将这些目标区分开来。我们将这种语义识别失败称为完全混叠。在一个受控的二元报告博弈中,在顺从上下文上拟合的真值探针与规定动作探针解决了相同的优化问题。在对抗上下文上,它们的标签互为补集,迫使它们的AUROC之和为1;这一恒等式在751个细胞层对上以浮点精度成立。我们使用随机码本将规定的输出符号与语义动作分离,然后通过在混合的顺从和对抗上下文上拟合,将真值与规定动作分离。对于一个经过奖励训练的Gemma-2-9B策略,其在所有评估的对抗试验中均给出虚假回答,传统探针在三个训练种子上的得分为$0.006 \pm 0.005$ AUROC,而混合拟合探针在相同的保留激活上得分为$1.000$。混合拟合使用了更多的训练样本并访问了带标签的对抗上下文,因此这一比较确立了线性可恢复性,而非孤立了解相关性的益处。我们还展示了两个顺从拟合探针,两者在分布内均表现完美,但在相同的对抗激活上得分分别为$0.080$和$0.986$。这些发现关乎探针所测量的内容:它们并未确立保留的功能信念、恢复方向的因果使用,或可部署的欺骗检测器。代码和汇总结果随论文附上。

英文摘要

Linear probes that decode the truth from a language model's activations have been proposed as deception monitors, and a truth probe scoring below chance on a model trained to deceive is naturally read as evidence that the model hid or stopped representing the truth. We show that a fitting-label ambiguity can produce the same readout. In a controlled game, a model should report a secret bit to an ally and its complement to a rival. On compliant (ally) contexts the true bit and the answer the task prescribes are identical labels, so a probe fitted there cannot tell which of the two it measures. We call this complete agreement perfect aliasing. The labels are complements on rival contexts, so one ally-fitted probe scored against each has rival AUROCs that sum to one. Mixed fitting, on ally and rival contexts together, makes the two labels differ; randomized output codebooks also decouple the prescribed answer from the output letter. For a reward-trained Gemma-2-9B policy that answers falsely on every evaluated rival trial, the ally-fitted probe scores $0.006 \pm 0.005$ AUROC at the final layer (mean $\pm$ sample SD over three RL training seeds), while mixed-fit probes score 1.000 on the same held-out activations. Mixed fitting also uses more examples and labelled rival data, so this shows the true bit remains linearly recoverable, not that separating the labels alone explains the gain. In instructed Llama-3.1-8B, a probe refitted on one prompt variant and one frozen from a reference variant both have held-out ally accuracy 1.000 but rival truth AUROCs of 0.080 and 0.986 on the same trials. Our main experiments use one single-token game that states the bit in the prompt, so the recovered direction may read a retained copy of it; where the model must infer the bit, no tested arm deceives reliably. We study what a probe measures, not whether the model uses that information.

发表机构

  • University of Waterloo(滑铁卢大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑