arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.06100cs.LGcs.AIcs.CL

VERPO:验证证据正则化策略优化

VERPO: Verified Evidence Regularized Policy Optimization

Haijiang Li, Chengyu Lv, Yi Zhang, Rui Qian, Zhibing Zhang, Xiangqing Shen, Junjie Yang, Yuchen Zhang, Wenyuan Jiang, Hanqing Hu, Cangqi Zhou

首次发表
浏览论文内容

中文总结 AI 辅助

VERPO框架利用证据修正策略,通过Fisher证据对比和ZPD控制器,在保留结果目标的同时提升科学推理和工具使用任务的平均得分。

中文摘要 AI 辅助

可验证的结果奖励指导语言模型的后训练,但序列级别的优势并不能识别哪些词元级别的决策应被保留或修改。证据条件教师通过重放带有特权反馈的采样轨迹来提供更密集的监督。然而,不加区分的模仿可能会转移不支持任务成功的格式或推理风格变化。我们引入了VERPO,一个验证证据正则化策略优化框架,将证据视为策略修正的提议,同时保留结果目标。它将无证据的参考恢复与有符号的词元级证据修正分开。Fisher证据对比沿着估计的证据存在方向衰减修正。一个停止的词元级ZPD控制器根据局部奖励对齐和Fisher移动成本来调整接受度,而参考通道独立于接受度。在五个科学推理和工具使用任务中,每个主干上的最佳变体在平均得分上超过了最强的比较基线。平均值在Qwen3-4B上从0.6826上升到0.6857,在Qwen3-8B上从0.6895上升到0.7058,在Llama-3.2-1B上从0.4751上升到0.5657。

英文摘要

Verifiable rewards improve language models through reliable task-level feedback, but methods based on Group Relative Policy Optimization (GRPO) apply a sequence-level advantage uniformly across all tokens. This coarse credit assignment reinforces or penalizes entire responses without identifying which local decisions to preserve, reinforce, or revise. Conversely, evidence-conditioned self-distillation provides denser token-level supervision, yet teacher imitation can transfer stylistic artifacts and miscalibrated confidence that destabilize training when misaligned with task success. We introduce VERPO, which converts evidence-conditioned guidance into reward-aligned token-level credit assignment while retaining the outcome objective. VERPO decomposes teacher guidance into an evidence-free reference term and signed, evidence-induced corrections at each token. A stopped controller combines selective acceptance, token-wise localization, and cost-aware scaling by balancing alignment with the local GRPO update direction against Fisher movement cost. Furthermore, we introduce Fisher Evidence Contrast (FEC), which attenuates nuisance shifts along an estimated evidence-presence direction through a regularized projection. Across five scientific reasoning and tool-use tasks, VERPO prevents optimization collapse and consistently achieves the highest multi-task average across model backbones, yielding marked improvements particularly on smaller models over strong baselines. Qualitative diagnostics confirm that token acceptance selectively targets reasoning bottlenecks consistent with local reward alignment and Fisher movement cost.

发表机构

  • Tongji University(同济大学)
  • Nanjing University(南京大学)
  • Fudan University(复旦大学)
  • Nanjing University of Science and Technology(南京理工大学)
  • Xiaohongshu(小红书)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑