arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CM-DPO:用于LLM规划的约束边际直接偏好优化

CM-DPO: Constraint-Margin Direct Preference Optimization for LLM Planning

Rabimba Karanjai, Qun Gu, Hemanth Hegadehalli Madhavarao, Wenhuan Sun, Xiaojiao Yu, Suryabhan Singh Hada, Libin N. George, Uma Kona, Richard Williamson, Linsey Pang, Prakhar Mehrotra

arXiv 2610.09219首次发表:更新:

发表机构

PayPal AI Lab(PayPal人工智能实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对DPO忽视约束严重性和偏好偏差的问题,提出CM-DPO,利用符号验证器生成连续边际并按严重程度缩放,分离硬软约束,在多个规划基准上8B模型达到89.2%通过率,匹配多智能体且延迟低13倍。

AI 中文摘要

直接偏好优化(DPO)将所有约束违反同等对待:1美元的预算超支和1000美元的超支产生相同的训练信号。当偏好对来自不同模型家族时,它也容易受到长度和风格偏差的影响。我们引入了约束边际DPO(CM-DPO),它用从确定性符号验证器导出的连续边际替换DPO的二元偏好信号,并按违反严重程度进行缩放。通过字典序目标分离硬约束和软约束,确保硬约束永远不会与偏好进行权衡。为了为CM-DPO提供偏差减少的训练对,我们通过程序化生成的约束配置文件(DCCG)和来自推理教师的最小编辑蒸馏(RT-MED)生成偏好数据,在一个我们称为SynPlan-R的框架内。在TravelPlanner、NaturalPlan和分布外PlanBench上,使用CM-DPO微调的8B模型实现了89.2%的通过率和93.4%的解决率,以13倍更低的延迟匹配多智能体系统,同时在未见过的Blocksworld上比GPT-4o高出9.2个百分点。

英文摘要

Direct Preference Optimization (DPO) treats all constraint violations equally: a $1 budget overshoot and a $1,000 overshoot induce the same training signal. It is also susceptible to length and style bias when preference pairs come from different model families. We introduce Constraint-Margin DPO (CM-DPO), which replaces DPO's binary preference signal with a continuous margin derived from a deterministic symbolic verifier and scaled by violation severity. Hard and soft constraints are separated through a lexicographic objective, ensuring hard constraints are never traded off against preferences. To supply CM-DPO with bias-reduced training pairs, we generate preference data through procedurally generated constraint profiles (DCCG) and minimal-edit distillation from a reasoning teacher (RT-MED), within a framework we call SynPlan-R. On TravelPlanner, NaturalPlan, and out-of-distribution PlanBench, an 8B model fine-tuned with CM-DPO achieves 89.2% pass rate and 93.4% solve rate, matching multi-agent systems at 13x lower latency while outperforming GPT-4o on unseen Blocksworld by 9.2 points.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑