通过优势解耦偏好优化实现视觉语言模型的统一生成与自验证
Unified Generation and Self-Verification for Vision-Language Models via Advantage Decoupled Preference Optimization
查看机构详情
- College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院)
- Venus Team, Ant Group(蚂蚁集团 Venus 团队)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
ADPO通过统一生成与自验证的强化学习框架,提升视觉语言模型的验证性能和推理效率。
中文摘要 AI 辅助
并行测试时间扩展通常训练单独的生成和验证模型,导致高训练和推理成本。我们提出优势解耦偏好优化(ADPO),一种联合学习答案生成和自验证的统一强化学习框架。ADPO引入了两个创新:偏好验证奖励提升验证能力以及解耦优化机制,使生成和验证能够协同优化。具体来说,偏好验证奖励通过正负样本计算平均验证分数作为决策阈值,在预测正确性与答案正确性一致时提供正反馈。同时,优势解耦优化计算生成和验证的独立优势,应用令牌掩码隔离梯度,并结合掩码GRPO目标,保留生成质量的同时校准验证分数。ADPO在验证AUC上达到最高+34.1%提升,在推理时间上降低53.5%,在MathVista/MMMU上获得显著提升+2.8%/+1.4%准确率,在ReasonSeg上获得+1.9 cIoU,在AndroidControl/GUI Odyssey上获得+1.7%/+1.0%步骤成功率。
英文摘要
Parallel test-time scaling typically trains separate generation and verification models, incurring high training and inference costs. We propose Advantage Decoupled Preference Optimization (ADPO), a unified reinforcement learning framework that jointly learns answer generation and self-verification within a single policy. ADPO introduces two innovations: a preference verification reward improving verification capability and a decoupled optimization mechanism enabling synergistic optimization of generation and verification. Specifically, the preference verification reward computes mean verification scores from positive and negative samples as decision thresholds, providing positive feedback when prediction correctness aligns with answer correctness. Meanwhile, the advantage decoupled optimization computes separate advantages for generation and verification, applies token masks to isolate gradients, and combines masked GRPO objectives, preserving generation quality while calibrating verification scores. ADPO achieves up to +34.1% higher verification AUC and -53.5% lower inference time, with significant gains of +2.8%/+1.4% accuracy on MathVista/MMMU, +1.9 cIoU on ReasonSeg, and +1.7%/+1.0% step success rate on AndroidControl/GUI Odyssey.