发表机构
Efficient Computation Inc.(高效计算公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过追踪探针准确率与干预效果,发现探针可解码上下文错误并引导修复,但可解码性无法证明输出信息丢失,最终状态优势未被检测到。
AI 中文摘要
先前的研究已证实,探针能够解码模型错误中的上下文绑定,且探针引导的干预可以修复其中部分错误。我们追踪了公开预训练和后训练检查点中的探针准确率、模型输出及干预响应。在Pythia预训练期间,探针准确率上升,而探针引导的干预从所有试验中可忽略的收益转变为在两种模型规模下更大的收益。保存的分数能够区分探针正确的错误,这些错误对正确候选具有低且高于均匀分布的模型概率。Oracle目标干预已能修复许多早期错误,但保存的聚合无法将目标质量与干预敏感性区分开来。对在最终状态或候选对数上训练的解码器进行的保留比较,在后期检查点模型错误上未发现最终状态的明显优势。一个信息论反例解释了为何仅凭错误上的可解码性无法确立被丢弃的输出信息。与下游遗漏的关联仍待探索。
英文摘要
Prior work established that a probe can decode an in-context binding on model errors and that probe-guided steering can repair some of them. We follow probe accuracy, model output, and steering response across public pretraining and post-training checkpoints. Probe accuracy rises during Pythia pretraining, while probe-guided steering moves from negligible all-trial benefit to a larger benefit at two model sizes. Saved scores distinguish probe-correct errors with low and above-uniform model probability for the correct candidate. Oracle-target steering already repairs many early errors, but saved aggregates cannot separate target quality from intervention sensitivity. A held-out comparison of decoders trained on the final state or candidate logits finds no detected final-state advantage on late-checkpoint model errors. An information-theoretic counterexample explains why decodability on errors alone cannot establish discarded output information. The connection to downstream omissions remains open.