AI 中文总结
针对多视图空间推理中思维链推理冗长不准确及标准GRPO奖励问题,提出LenGuard-GPC框架,通过比较标准与引导提示下预测分布得KL散度作奖励信号,并引入阶段长度奖励,提升了准确率且缩短了平均响应长度。
AI 中文摘要
多视图空间推理要求视觉语言模型跨图像比较视觉证据、对齐对象对应关系并在长视觉上下文中推断空间关系,在这种情况下,思维链推理往往会变得冗长而不准确。具有可验证奖励的强化学习适合此任务,但标准GRPO奖励依赖稀疏结果级反馈,无法提供推理轨迹错误位置的信号,也无法控制其长度。我们提出LenGuard-GPC,一个密集奖励框架,可同时解决这两个问题。对于每个采样轨迹,它比较标准提示和引导提示下的逐令牌预测分布,并使用由此产生的令牌总和KL散度作为密集奖励信号。由于此KL惩罚会累积在令牌上,否则会奖励较短的响应而不管其质量如何,因此我们引入了一个阶段长度奖励,将推理长度保持在可控范围内,而不是简单地鼓励简短。在六个多视图空间推理基准上,LenGuard-GPC提高了准确率,同时减少了平均响应长度。
英文摘要
Multi-view spatial reasoning requires vision-language models to compare visual evidence across images, align object correspondences, and infer spatial relations over long visual contexts, a setting where chain-of-thought reasoning tends to grow verbose without becoming more accurate. Reinforcement learning with verifiable rewards is a natural fit for this task, but standard GRPO reward relies on sparse outcome-level feedback and gives no signal about where a reasoning trajectory goes wrong, nor any control over its length. We propose LenGuard-GPC, a dense reward framework that addresses both problems together. For each sampled trajectory, it compares the token-wise predictive distributions under a standard prompt and a guided prompt, and uses the resulting token-sum KL divergence as a dense reward signal. Since this KL penalty accumulates over tokens and would otherwise reward shorter responses regardless of their quality, we introduce a staged length bonus that keeps reasoning length within a controlled range without simply encouraging brevity. On six multi-view spatial reasoning benchmarks, LenGuard-GPC improves accuracy over vanilla GRPO while reducing average response length.