发表机构
Facebook; University of California, Irvine; Harbin Institute of Technology; FIS; ByteDance Inc(Facebook; 加利福尼亚大学尔湾分校; 哈尔滨工业大学; FIS; 字节跳动公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对多模态推理中自然语言CoT开销大、视觉基础弱的问题,提出DRT范式,用紧凑结构化轨迹表达推理,通过轨迹基础强化学习提升token效率5.5倍并提高准确率1.3点。
AI 中文摘要
尽管多模态大语言模型(MLLMs)取得了显著进展,但主流的思维链(CoT)范式仍局限于自然语言表达空间。因此,它们固有地产生过多的语言开销,导致信息稀释和视觉基础薄弱。为了解决这一挑战,我们提出了稠密推理轨迹(DRT),这是一种偏离以自然语言为中心的CoT的范式,通过将推理表达为紧凑的结构化轨迹,包括带有符号连接器的简洁中间状态,并将视觉观察与逻辑推理分离。首先,我们引入稠密轨迹初始化,将DRT推理模式内化到模型中,在保留视觉证据的同时显著提高token效率。为了进一步使模型忠实捕获轨迹内的逻辑关系,我们提出了轨迹基础强化学习框架,该框架通过三视角验证流程构建参考轨迹,并采用带有结构化奖励的轨迹基础GRPO,鼓励模型生成简洁的DRT风格轨迹,减少幻觉并增强逻辑基础。在具有挑战性的推理基准上的大量实验表明,与Qwen3-VL基线相比,DRT实现了5.5倍的token效率提升,同时提高了1.3个准确率点。这些发现表明,复杂的多模态推理可能不需要冗长的自然语言轨迹,为下一代MLLMs开辟了一条更高效的路径。我们的代码和数据可在以下网址获取:this https URL
英文摘要
Despite the remarkable progress in Multimodal Large Language Models (MLLMs), prevailing Chain-of-Thought (CoT) paradigms remain confined to the natural-language expression space. Consequently, they inherently incur excessive linguistic overhead, leading to information dilution and weak visual grounding. To address this challenge, we propose Dense Reasoning Trace (DRT), a paradigm that departs from natural-language-centered CoT by expressing reasoning as compact structured traces, which include concise intermediate states with symbolic connectors and disentangle visual observations from logical deductions. First, we introduce the Dense Trace Initialization to internalize the DRT reasoning mode into the model, substantially improving token efficiency while preserving visual evidence. To further enable the model to faithfully capture the logical relations within traces, we propose the Trace-Grounded Reinforcement Learning framework, which builds reference traces through a tri-perspective verification pipeline and employs Trace-Grounded GRPO with structured rewards, encouraging the model to generate concise DRT-style traces with reduced hallucination and stronger logical grounding. Extensive experiments on challenging reasoning benchmarks show that DRT achieves 5.5$\times$ token efficiency improvement while improving 1.3 accuracy points over the Qwen3-VL baseline. These findings suggest that complex multimodal reasoning may not require verbose natural-language traces, opening a more efficient path for next-generation MLLMs. Our code and data are available at: https://github.com/HIT-leaderone/DRT