arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

双向偏好合成:从边界失败中学习提示条件偏好

Bidirectional Preference Synthesis: Learning Prompt-Conditioned Preferences from Boundary Failures

Junbo Wang, Lidong Lu, Zhuoqun Li, Guiping Jiang, Xiangyu Wu, Tinghai Zhang, Tong Lu

arXiv 2610.04328首次发表:更新:

发表机构

Kuaishou Technology; Nanjing University(快手科技; 南京大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对边界失败中提示依赖缺失的问题,提出双向偏好合成(BPS)方法,通过为每个失败样本构建正反向配对,在不改变DPO目标下将已实现侧排名准确率从6.8%提升至62.3%,并保持能力保留。

AI 中文摘要

基于修正的离线偏好流水线通常仅将模型失败视为原始提示下的被拒绝响应。这种监督对于边界失败是不完整的:边界失败是指违反给定指令但连贯地满足附近意图或约束设置的响应。我们引入了双向偏好合成(BPS),这是一种用于标准直接偏好优化(DPO)的数据构建方法,它使这种缺失的提示依赖性变得明确。对于每个经过验证的边界失败,BPS在原始提示下保留传统的正向配对,并在合成的已实现提示下添加反向配对,因此同一响应在其错误之处被拒绝,在其正确之处被选择,而无需改变DPO目标、训练奖励模型或要求在线采样。在Qwen3-4B-Instruct-2507上,BPS保持了原始侧成对排名,同时将保留的交叉锚点上的已实现侧排名准确率从6.8%提高到62.3%,在Kimi-K2.6跨教师探测下也有类似的转变。一项盲人人工审计支持预期的反向偏好方向,下游评估显示,在多语言多轮指令遵循中,与正向DPO的分离最为清晰,在智能体、工具使用和代码检查上具有一致的能力保留模式。

英文摘要

Correction-based offline preference pipelines commonly treat model failures only as rejected responses under the original prompt. This supervision is incomplete for boundary failures: responses that violate the given instruction yet coherently satisfy a nearby intent or constraint setting. We introduce Bidirectional Preference Synthesis (BPS), a data-construction method for standard Direct Preference Optimization (DPO) that makes this missing prompt dependence explicit. For each validated boundary failure, BPS keeps the conventional forward pair under the original prompt and adds a reverse pair under a synthesized achieved prompt, so the same response is rejected where it is wrong and chosen where it is right, without changing the DPO objective, training a reward model, or requiring online sampling. On Qwen3-4B-Instruct-2507, BPS preserves original-side pairwise ranking while raising achieved-side ranking accuracy from 6.8% to 62.3% on held-out crossed anchors, with a similar shift under a Kimi-K2.6 cross-teacher probe. A blind human audit supports the intended reverse preference direction, and downstream evaluations show the clearest separation from Forward-DPO in multilingual multi-turn instruction following, with consistent capability-retention patterns on agentic, tool-use, and code checks.

Comments18 pages, 3 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑