arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.10889cs.CV

SPLIT-RL:基于声明级优势的分阶段感知-语言推理训练

SPLIT-RL: Staged Perception-Language Reasoning Training with Claim-Level Advantages

Raja Kumar, Rajat Koner, Ritwick Chaudhry, Zhuowei Li, Nishant Sankaran, Yifan Xing

首次发表
浏览论文内容

中文总结 AI 辅助

SPLIT-RL是一种分阶段训练视觉推理与语言推理的强化学习方法,引入声明级优势CLA-GRPO,在多类VL模型上较GRPO提升1.4-6.1个百分点,同时改善视觉与语言推理能力。

中文摘要 AI 辅助

视觉-语言(VL)推理要求模型既能从图像中提取相关且准确的信息(视觉推理,VR),又能据此推断出答案(语言推理,LR)。带有可验证奖励的强化学习通常通过单一思维链(CoT)结合最终答案奖励来同时训练这两部分,这使得每个CoT token都具有相同的序列级优势,无法区分特定能力的错误。我们提出SPLIT-RL,一种分阶段的后训练方法,在不重叠的阶段中分别训练VR和LR。由于每组的rollout仅在一种能力上存在差异,组相对优势可将其分离,且每个阶段使用阶段特定的奖励进行优化。我们进一步引入声明级优势(CLA-GRPO),将VR阶段的rollout分解为原子视觉声明,并基于视觉类型的组形成提供声明级的细粒度优势。尽管分两阶段训练,训练后的策略在推理时仍像GRPO模型一样,通过单次CoT调用进行评估。在该协议下,SPLIT-RL在2B至30B-A3B的Qwen3-VL模型及InternVL3.5-8B上,平均准确率较GRPO提升1.4至6.1个百分点。使用基于神谕的诊断评估每种能力显示,仅使用答案的GRPO未改变感知能力,而SPLIT-RL同时提升了VR和LR能力。

英文摘要

Vision-Language (VL) reasoning requires a model to both extract relevant and accurate information from an image (visual reasoning, VR), and to infer the answer from it (language reasoning, LR). Reinforcement learning with verifiable rewards typically trains both through a single chain-of-thought with a final-answer reward. This gives every CoT token the same sequence-level advantage, failing to distinguish capability specific errors. We propose SPLIT-RL, a staged post-training approach that trains VR and LR in disjoint phases. Because a group's rollouts differ along one capability at a time, the group-relative advantage isolates it, and each phase is optimized using phase-specific reward. We further introduce Claim-Level Advantage (CLA-GRPO), which decomposes VR-phase rollouts into atomic visual claims and provides a fine-grained advantage at claim level based on visual-type group formation. Although trained in two phases, trained policy is evaluated like GRPO model, with a single CoT call at inference time. Under this protocol, SPLIT-RL improves average accuracy over GRPO by 1.4-6.1 points across Qwen3-VL models from 2B to 30B-A3B and InternVL3.5-8B. Evaluating each capability using an oracle based diagnostic shows that answer-only GRPO leaves perception unchanged, whereas SPLIT-RL improves both VR and LR.

发表机构

  • University of Southern California(南加州大学)
  • Amazon AGI(亚马逊AGI)

机构由 AI 辅助整理,请以论文原文为准。

↑