AI 中文总结
通过偏好优化在可靠性标记数据上训练,模型学会根据来源可靠性调整行为,实现依赖先验的可靠性开关,提升决策准确性。
AI 中文摘要
一个被对齐的模型,在被要求面对操纵性来源时坚持自己的答案,但仍然必须对可靠的来源进行更新,并拒绝不可靠的来源:抵抗、可靠更新和不可靠来源拒绝是一个三方契约,而不是三个独立的行为。我们表明,大多数反谄媚工作所优化的目标在来源可靠性方面是不可识别的:因为没有任何偏好标签依赖于来源是否实际可靠,任何臂的标量混合都追踪一个单一的顺从刻度,并且没有任何点能区分两个仅陈述可靠性不同的同模板证词。这种固定↔轻信前沿是目标的一个属性,而不是任何模型的属性。我们通过数据使可靠性可识别:一个阈值基准,其中来源在陈述其可靠性r的同时断言相反的答案,正确的行动是当且仅当r超过模型的先验强度p时翻转。在平衡覆盖上的偏好优化安装了一个依赖先验的可靠性开关:在Qwen2.5-7B-Instruct上的三个种子中,阈值r*随先验单调上升,决策准确率达到0.84,具有单调翻转曲线(Spearman 0.56),并且策略泛化到未见过的可靠性值和保留的符号,遵循陈述的可靠性而非角色声望。三个对照实验定位了原因:一个不匹配的变体同样安装开关(0.80),第二个偏好优化器(IPO)同样安装得好(0.86),而监督模仿则不能(0.50),因此原因是偏好优化在可靠性标记的覆盖上,而不是配对、损失或模仿。一个确认性电池在新鲜的测试抽取上复制了开关,诚实地界定它(它基于证词中陈述的可靠性,而不是单独审计的记录),并将其转移到Llama-3.1-8B。该前沿是经验性的,而不是一个定理。
英文摘要
An aligned model asked to hold its answer against a manipulative source must still update on a reliable one and reject an unreliable one: resistance, reliable-update, and unreliable-source rejection are one three-way contract, not three independent behaviors. We show the objective most anti-sycophancy work optimizes is non-identifying with respect to source reliability: because no preference label depends on whether a source is actually reliable, any scalar mixture of the arms traces a single deference dial, and no point separates two same-template testimonies differing only in stated reliability. This fixation$\leftrightarrow$gullibility frontier is a property of the objective, not any model. We make reliability identifiable through data: a threshold benchmark where a source asserts the opposite answer while stating its reliability $r$, and the correct action is to flip iff $r$ exceeds the model's prior strength $p$. Preference optimization over balanced coverage installs a prior-dependent reliability switch: across three seeds on Qwen2.5-7B-Instruct the threshold $r^\star$ rises monotonically with the prior, decision accuracy reaches $0.84$ with a monotone flip curve (Spearman $0.56$), and the policy generalizes to unseen reliability values and a held-out notation, following stated reliability over role prestige. Three controls localize the cause: an unmatched variant installs the switch equally ($0.80$), a second preference optimizer (IPO) installs it just as well ($0.86$), whereas supervised imitation does not ($0.50$), so the cause is preference optimization over reliability-labeled coverage, not pairing, loss, or imitation. A confirmatory battery replicates the switch on a fresh test draw, bounds it honestly (it keys on reliability stated in the testimony, not a separately audited record), and transfers it to Llama-3.1-8B. The frontier is empirical, not a theorem.