arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.13849cs.AI

UniCAR-RL:在视觉数学中先看得更清,再想得更深

UniCAR-RL: Seeing Better before Thinking Deeper in Visual Mathematics

Yuzhe Li, Hao Yan, Hao Wang, Xingchen Liu, Ya-Qi Yu, Jihao Wu, Minghui Liao, Wei Chen, Yuliang Liu

首次发表
浏览论文内容

中文总结 AI 辅助

针对多模态大模型视觉数学推理中感知幻觉引发级联错误的问题,提出UniCAR-RL,通过解耦感知与推理优化,以三个协同分支提升推理性能,仅用短答案数据即显著增强并具良好泛化性。

中文摘要 AI 辅助

多模态大语言模型(MLLMs)在复杂的数学视觉推理中常常表现不佳,主要原因是缺乏细粒度的感知能力,导致初始的视觉幻觉直接引发级联的推理失败。在传统的端到端强化学习(RL)中,稀疏奖励无法将感知幻觉与逻辑失误解耦,阻碍了针对性的感知优化。另外,使用感知增强的思维链(CoT)数据进行微调成本高昂且易产生幻觉。在本文中,我们通过提出UniCAR-RL——一个无需标注的强化学习框架——来解决这些挑战。通过在训练过程中显式解耦感知与推理的优化,该框架实现了两种能力的隔离与优化。具体而言,UniCAR-RL由三个协同分支组成:1)Caption-RL分支,通过验证器引导的推理验证来优化感知能力;2)Reasoning-RL分支,基于黄金图像描述进行逻辑推理以阻断级联错误;3)QA-RL分支,保留原生端到端对齐以确保稳健的问答性能。实验表明,UniCAR-RL仅使用原始的短答案数据即可显著提升MLLMs的数学与视觉推理能力。此外,它在多种架构和规模上展现出强大的泛化能力。

英文摘要

Multimodal Large Language Models (MLLMs) often struggle with complex mathematical visual reasoning primarily due to a lack of fine-grained perception, causing initial visual hallucinations to directly trigger cascading reasoning failures. In traditional end-to-end reinforcement learning (RL), sparse rewards fail to decouple perceptual hallucinations from logical missteps, hindering targeted perception optimization. Alternatively, fine-tuning with perception-enhanced CoT data incurs high costs and hallucinations. In this paper, we address these challenges by proposing UniCAR-RL, an annotation-free RL framework. By explicitly decoupling the optimization of perception and reasoning during the training process, it achieves isolation and optimization of both capabilities. Specifically, UniCAR-RL consists of three synergistic branches: 1) a Caption-RL branch that optimizes perception capabilities through verifier-guided reasoning validation; 2) a Reasoning-RL branch that performs logical reasoning based on a gold image description to halt cascading errors; 3) a QA-RL branch that retains native end-to-end alignment to ensure robust question-answering performance. Experiments show that UniCAR-RL substantially improves MLLMs' mathematical and visual reasoning using only raw short-answer data. Furthermore, it demonstrates strong generalization across diverse architectures and scales.

发表机构

  • Huawei Inc.(华为公司)
  • Huazhong University of Science and Technology(华中科技大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑