发表机构
PayPal AI Lab(PayPal人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对DPO忽视约束严重性和偏好偏差的问题,提出CM-DPO,利用符号验证器生成连续边际并按严重程度缩放,分离硬软约束,在多个规划基准上8B模型达到89.2%通过率,匹配多智能体且延迟低13倍。
AI 中文摘要
直接偏好优化(DPO)将所有约束违反同等对待:1美元的预算超支和1000美元的超支产生相同的训练信号。当偏好对来自不同模型家族时,它也容易受到长度和风格偏差的影响。我们引入了约束边际DPO(CM-DPO),它用从确定性符号验证器导出的连续边际替换DPO的二元偏好信号,并按违反严重程度进行缩放。通过字典序目标分离硬约束和软约束,确保硬约束永远不会与偏好进行权衡。为了为CM-DPO提供偏差减少的训练对,我们通过程序化生成的约束配置文件(DCCG)和来自推理教师的最小编辑蒸馏(RT-MED)生成偏好数据,在一个我们称为SynPlan-R的框架内。在TravelPlanner、NaturalPlan和分布外PlanBench上,使用CM-DPO微调的8B模型实现了89.2%的通过率和93.4%的解决率,以13倍更低的延迟匹配多智能体系统,同时在未见过的Blocksworld上比GPT-4o高出9.2个百分点。
英文摘要
Direct Preference Optimization (DPO) treats all constraint violations equally: a $1 budget overshoot and a $1,000 overshoot induce the same training signal. It is also susceptible to length and style bias when preference pairs come from different model families. We introduce Constraint-Margin DPO (CM-DPO), which replaces DPO's binary preference signal with a continuous margin derived from a deterministic symbolic verifier and scaled by violation severity. Hard and soft constraints are separated through a lexicographic objective, ensuring hard constraints are never traded off against preferences. To supply CM-DPO with bias-reduced training pairs, we generate preference data through procedurally generated constraint profiles (DCCG) and minimal-edit distillation from a reasoning teacher (RT-MED), within a framework we call SynPlan-R. On TravelPlanner, NaturalPlan, and out-of-distribution PlanBench, an 8B model fine-tuned with CM-DPO achieves 89.2% pass rate and 93.4% solve rate, matching multi-agent systems at 13x lower latency while outperforming GPT-4o on unseen Blocksworld by 9.2 points.