arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

抵抗与更新:用于激励兼容大语言模型的反事实报告坐标

Resist and Update: Counterfactual Report Coordinates for Incentive-Compatible LLMs

Sen Yang, Yuen-Hei Yeung

arXiv 2607.12985首次发表:更新:

AI 中文总结

研究大语言模型在非证据激励压力下误报问题,提出学习和验证反事实报告调解器方法,通过互换干预识别低秩报告坐标,引入CRC钳位,在基准测试中取得成果,贡献了接口和认证方法。

AI 中文摘要

对齐的语言模型在非证据激励压力下经常误报:即使内部信念不变,它们也会认同自信的用户或夸大确定性。我们将此视为内部激励兼容性(IC)的失败,并提出一种学习和验证反事实报告调解器的方法,该调解器使模型的报告符合因果契约:对禁止的影响(压力、声誉、重新措辞)不变,对许可的影响(真实证据)有响应。这两个要求,即抵抗和更新,方向相反。我们在具有已知后验的贝叶斯见证基准上对它们进行研究,在该基准中,相同的用户分歧纯粹根据所述来源可靠性是许可证据或禁止压力。我们(i)通过互换干预而非探测准确性因果识别答案、置信度和警告的低秩报告坐标,这些坐标近乎正交且可独立控制,(ii)引入一种无需训练的反事实报告坐标(CRC)钳位,该钳位在反事实激励中和的背景下参考模型自己的报告。在见证基准上,双程钳位联合实现了1.00的抵抗和更新(威尔逊95%置信区间[0.99,1.00]),这是在可构建参考下的因果证书,而非部署的解决方案。全局解码和引导显示单参数权衡;仅当两者都被枚举时,输出级微调才匹配两个目标;仅抵抗训练会失去证据响应性。可部署的单程编译有损失(0.73/0.97)。该机制和钳位在三个模型家族中重现并转移到自然谄媚基准(SycophancyEval)。我们的贡献是接口和认证方法:激活级反事实激励不变性作为内部IC的结构原语。

英文摘要

Aligned language models routinely misreport under non-evidential pressure: they cave to a confident user, yet fail to revise when genuine evidence arrives. We cast this as a failure of internal incentive-compatibility and study the two demands, resist (ignore forbidden pressure) and update (follow licensed evidence), on a Bayesian-witness benchmark with known posteriors, where the same user disagreement is evidence or pressure purely by stated source reliability, removing the evidence/pressure confound by construction. Using interchange interventions rather than probes, we causally localize low-rank report coordinates for answer, confidence, and caveat, establishing causal sufficiency at a late intervention site rather than uniqueness or necessity, with a causal cross-talk matrix showing strong own-coordinate control and only small cross-effects (partial functional disentanglement). We then introduce a training-free counterfactual report-coordinate (CRC) clamp that references the model's own report under an incentive-neutralized counterfactual of the prompt. The two-pass full-window clamp attains resist and update of $1.00$ jointly (Wilson 95% CI $[0.99,1.00]$; the rank-16 projection alone reaches $0.88/0.90$), which we read as a causal certificate and upper bound under a constructible reference, not a claim of a deployed solution. Tested global decoding and fixed-direction steering trade one objective against the other, and resist-only training collapses updating to $0.01$. The deployable single-pass compilation is lossy ($0.73/0.97$). The mechanism and the clamp reproduce across three model families and transfer to a natural sycophancy benchmark with significant paired improvements. Our contribution is the interface and certification method: activation-level counterfactual incentive-invariance as a structural primitive for internal incentive-compatibility.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑