发表机构
Tsinghua University; Microsoft Research(清华大学; 微软研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对多模态几何推理中现有方法的缺陷,提出信用可寻址推理,通过Code-CoT和CE-GRPO实现,在9个几何基准上平均准确率达76.04,优于对比模型,凸显协同设计的价值。
AI 中文摘要
多模态几何推理要求视觉语言模型(VLMs)提取精确的视觉关系,并通过多步骤推理保留这些关系。现有的自由形式轨迹会模糊决定答案的决策,而轨迹级强化学习会将单个终端信号分布到整个响应中。我们提出信用可寻址推理,其中推理过程中暴露的语义单元同时定义了学习时比较备选方案和分配信用的位置。我们通过Code-CoT实例化该原理,它保留了图表,将视觉关系表示为行可寻址可执行代码,并将推理组织为类型化事件;还通过CE-GRPO实例化,它使用结构先验和类型归一化熵选择事件边界,从共享前缀中采样完整的延续,并将结果差异转换为局部优势。在9个几何基准测试中,CE-GRPO的平均准确率达到76.04,分别优于Qwen3-VL-8B和轨迹级GRPO 8.09和3.43个百分点。其相对优势随中间事件数量的增加而提升,证明了表示-优化协同设计对于长且依赖密集的多模态推理的价值。
英文摘要
Multimodal geometry reasoning requires VLMs to extract precise visual relations and preserve them through multi-step deduction. Existing free-form traces obscure the decisions that determine the answer, and trajectory-level reinforcement learning distributes a single terminal signal across the entire response. We introduce credit-addressable reasoning, in which the semantic units exposed during inference also define where learning compares alternatives and assigns credit. We instantiate this principle with Code-CoT, which retains the diagram, represents visual relations as line-addressable executable code, and organizes reasoning into typed events, and CE-GRPO, which selects event boundaries using structural priors and type-normalized entropy, samples complete continuations from shared prefixes, and converts outcome differences into localized advantages. Across nine geometry benchmarks, CE-GRPO achieves an average accuracy of 76.04, outperforming Qwen3-VL-8B and trajectory-level GRPO by $8.09$ and 3.43 points, respectively. Its relative advantage increases with the number of intermediate events, demonstrating the value of representation--optimization co-design for long, dependency-heavy multimodal reasoning.