发表机构
Istanbul Technical University; Pabna University of Science and Technology; American International University-Bangladesh; Deakin University; North South University(伊斯坦布尔理工大学; 帕布纳科技大学; 美国国际大学-孟加拉; 迪肯大学; 南北大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对思维链忠实性检测器在分布偏移下自身不稳定的问题,提出元忠实性原理及SIFT检测器,证明主要瓶颈是采样随机性而非偏移,SIFT可减少64%的不变性违反。
AI 中文摘要
思维链(CoT)忠实性检测器被广泛用于审计推理模型,然而检测器本身也是一个预测器,其判定结果被视为稳定属性。我们提出一个问题:在分布偏移下,检测器对自身是否忠实?我们将元忠实性形式化为一个不变性原理:一个有效的检测器必须在仅由保持真实忠实性的变换所区分的轨迹上返回相同的判定。我们证明了三个结果:(i)仅使用干预-响应轮廓的检测器无法区分具有相同签名的忠实机制与附带现象机制;(ii)任何依赖偏移敏感特征的检测器,其违反不变性的比率与其分布内准确率无关;(iii)渐近认证的选择性风险保证可实现自信的弃权(不执行)。我们将该原理应用于FaithShift——一个涵盖十个偏移轴的压力测试协议,并提出了SIFT——一种使用跨环境不变性目标和认证弃权训练的隐藏状态轨迹检测器。在14,996条轨迹、四个领域和八个模型上,出现了三个发现。第一,迁移崩溃是真实存在的:所有现有检测器的差距均≥0.15 AUROC。第二,主要瓶颈是采样随机性而非偏移:超过80%的检测器不稳定性源于随机种子变化,这否定了我们预先注册的预测,即偏移导致的违反率超过0.25。第三,SIFT将不变性违反率比最佳单种子基线降低了64%,但任何检测器的四种子集成将差距缩小到0.01(在匹配覆盖率下无显著差异,p=0.21),且SIFT需要51%的弃权率。跨模型迁移从族内到族间再到开放权重到API逐渐退化,部分通过多模型训练得以弥补。我们提供了一个审计审计者的框架:真正的障碍是检测器方差,而非分布偏移。
英文摘要
Chain-of-Thought (CoT) faithfulness detectors are widely used to audit reasoning models, yet a detector is itself a predictor whose verdicts are treated as stable properties. We ask whether a detector is faithful to itself under distribution shift. We formalize meta-faithfulness as an invariance principle: a valid detector must return identical verdicts on traces that differ only by transformations preserving ground-truth faithfulness. We prove three results: (i) no detector using only intervention-response profiles can separate faithful from epiphenomenal mechanisms with identical signatures; (ii) any detector relying on shift-sensitive features violates invariance at a rate independent of its in-distribution accuracy; (iii) an asymptotic certified selective-risk guarantee enables confident abstention. We operationalize the principle in FaithShift, a stress-test protocol spanning ten shift axes, and propose SIFT, a hidden-state trajectory detector trained with cross-environment invariance objectives and certified abstention. Across 14,996 traces, four domains, and eight models, three findings emerge. First, transfer collapse is real: all existing detectors show gaps $\geq 0.15$ AUROC. Second, the dominant bottleneck is sampling stochasticity, not shift: over 80% of detector instability stems from random seed variation, falsifying our preregistered prediction that shift-attributable violations exceed 0.25. Third, SIFT cuts invariance violations by 64% over the best single-seed baseline, but a four-seed ensemble of any detector narrows the margin to 0.01 (indistinguishable at matched coverage, $p=0.21$), and SIFT needs a 51% abstention rate. Cross-model transfer degrades from within-family to cross-family to open-weight-to-API, partly closed by multi-model training. We offer a framework for auditing auditors: the real barrier is detector variance, not distribution shift.
CommentsUnder review as a conference paper at ICLR 2027