arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过选择性上下文偏好优化学习何时信任

Learning When to Trust via Selective Context Preference Optimization

Xian Sun, Wei Chow, Yingshuo Wang, Junhao Liu, Wei Gao, Qing Wu, Lingdong Kong

arXiv 2608.06377首次发表:更新:

发表机构

Duke University; National University of Singapore; UC Berkeley; UC Irvine; Northeastern University; Nanyang Technological University(杜克大学; 新加坡国立大学; 加州大学伯克利分校; 加州大学欧文分校; 东北大学; 南洋理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对语言模型易受误导性外部信号影响的问题,提出SCOPE方法,在含四种条件的MIST基准上优化DPO目标,降低SC2W的同时保持正常上下文下的准确率,主张以选择性信任评判模型。

AI 中文摘要

语言模型越来越多地基于外部信号生成答案,而单个误导性信号就能将正确答案变为错误答案。明显的补救措施是训练模型抵制此类信号,但这隐藏了一种失效模式:忽略所有上下文的模型看似鲁棒,却在上下文值得信任时毫无用处。我们将此问题重新定义为选择性信任,并推出MIST,这是一个人工标注的基准,为每个推理项提供四种匹配条件(干净、误导性、正确上下文、不相关上下文),同时提出SC2W,这是一种配对指标,用于统计误导性信号将干净正确答案转为错误答案的频率。在全面的基准研究中,我们发现这种易受影响性是普遍存在的。随后我们提出SCOPE,该方法挖掘干净正确/误导错误的失效案例,并在所有四种条件下均衡的匹配偏好对上优化标准直接偏好优化(DPO)目标,而非仅针对误导性项。我们的方法大幅降低了流行开源模型的SC2W,同时在添加的上下文干净、正确或不相关时保持准确率。通过这项工作,我们认为应基于选择性信任而非仅抵制来评判模型。

英文摘要

Language models increasingly condition their answers on external signals, and a single misleading one can turn a correct answer wrong. The obvious remedy, training models to resist such signals, hides a failure mode: a model that ignores all context looks robust yet is useless when the context is worth trusting. We recast the problem as selective trust and introduce MIST, a human-annotated benchmark that renders each reasoning item under four matched conditions (clean, misleading, correct-context, and irrelevant-context), together with SC2W, a paired metric counting how often a misleading signal flips a clean-correct answer to wrong. Across a comprehensive benchmark study, we observe that such a susceptibility is universal. We then propose SCOPE, which mines clean-correct/misleading-wrong failures and optimizes a standard Direct Preference Optimization (DPO) objective over matched preference pairs balanced equally across all four conditions, rather than over misleading items alone. Our approach substantially reduces SC2W on popular open-sourced models while preserving accuracy when the added context is clean, correct, or irrelevant. With this work, we argue that models should be judged on selective trust, not on resistance alone.

CommentsProject Page at https://worldbench.github.io/scope GitHub Repo at https://github.com/worldbench/SCOPE HF Dataset at https://huggingface.co/datasets/worldbench/MIST-Bench

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑