arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

代码一致性偏好优化验证用于语言模型对齐

Code Consistency Preference Optimization Verification for Language Model Alignment

Yunlong Tan, Mingqiao Mo, Hao Zhang

arXiv 2609.19002首次发表:更新:

发表机构

University of Chinese Academy of Sciences(中国科学院大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出基于依赖图和执行一致性分数的偏好优化方法,微调Llama-3-8B和DeepSeekMath-7B,在数学推理和物理推理上超越现有模型,形成CCPO模型系列。

AI 中文摘要

基于执行的验证通过计算健全性和依赖感知过滤增强了大型语言模型的数学推理能力。然而,先前依赖Bradley-Terry奖励模型的偏好优化方法未能捕捉科学任务所需的逻辑依赖和执行一致性。我们提出了一种方法,通过依赖图生成计算上健全的解决方案,用于执行一致的偏好优化。我们首先使用UltraFeedback提示、模型生成、验证和一致性结果构建了一个科学推理数据集。然后,我们提取推理步骤表达式、前提条件和可推导性关系,以构建依赖图并计算执行一致性分数。这些分数被附加到每个步骤,创建配对训练数据。微调Llama-3-8B和DeepSeekMath-7B取得了显著提升:在MATH上提高了17.0%,在GSM8K上提高了15.1%。扩展我们的科学可行性控制框架在PhyX多模态物理推理上达到了50.1%的准确率,超过了DeepSeek-R1(49.8%)和OpenAI o3-mini(48.2%),在alpha=0.10时科学有效性覆盖率达到91.7%,科学定律违规减少了73%,最终形成了CCPO模型系列。

英文摘要

Execution-based verification enhances large language models' mathematical reasoning through computational soundness and dependency-aware filtering. However, prior preference optimization methods relying on Bradley-Terry reward models fail to capture the logical dependencies and execution consistency needed for scientific tasks. We propose a method that generates computationally sound solutions with dependency graphs for execution-consistent preference optimization. We first build a scientific reasoning dataset using UltraFeedback prompts, model generations, verification, and consistency results. Then we extract reasoning step expressions, prerequisites, and derivability relationships to construct dependency graphs and compute execution consistency scores. These scores are appended to each step, creating paired training data. Fine-tuning Llama-3-8B and DeepSeekMath-7B yields significant gains: +17.0% on MATH and +15.1% on GSM8K. Extending our Scientific Feasibility Control framework achieves 50.1% accuracy on PhyX multimodal physics reasoning, surpassing DeepSeek-R1 (49.8%) and OpenAI o3-mini (48.2%), with 91.7% scientific validity coverage at alpha=0.10 and 73% fewer scientific law violations, resulting in the CCPO model family.

Comments24 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑