发表机构
Beijing University of Posts and Telecommunications; BUPT Shenzhen Institute; Carnegie Mellon University; Lancaster University; Nanyang Technological University; China Telecom Corporation Limited Sichuan Branch; Changsha University of Science & Technology(北京邮电大学; 北京邮电大学深圳研究院; 卡内基梅隆大学; 兰卡斯特大学; 南洋理工大学; 中国电信股份有限公司四川分公司; 长沙理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究多模态推理中模型中间步骤有缺陷的问题,提出$O^2$-CritiCuRL框架,通过离线多步展开分析和在线渐进式强化学习策略,区分关键与冗余步骤,提升模型性能及训练和推理效率。
AI 中文摘要
多模态大语言模型在推理任务中展现出能力,但常产生有缺陷的中间步骤,影响可解释性和可靠性。尽管已有对步骤级监督的探索,但区分决定性步骤和冗余步骤仍具挑战。我们提出了$O^2$-CritiCuRL,一种新颖的课程强化学习框架,通过迭代离线-在线范式引入关键步骤意识。离线阶段通过对带步骤注释的轨迹进行多步展开分析估计步骤重要性,过滤冗余步骤;在线阶段采用渐进式步骤级强化学习策略,由截断链引导模型推断缺失步骤并完善推理。在多模态推理基准上的大量实验表明,该方法取得了最优性能,同时具有更高的训练和推理效率。
英文摘要
Multimodal large language models exhibit capabilities on reasoning tasks, yet often produce flawed intermediate steps while yielding correct final answers. This behavior undermines interpretability and reliability, suggesting reliance on spurious shortcuts rather than faithful reasoning. Although efforts have explored step-level supervision, distinguishing decisive steps from redundant ones remains challenging. We propose $O^2$-CritiCuRL, a novel curriculum reinforcement learning framework that introduces critical-step awareness through an iterative offline-online paradigm. In the offline stage, $O^2$-CritiCuRL conducts multi-rollout analysis over step-annotated trajectories to estimate step-level importance, allowing the framework to distill critical reasoning steps and filter out redundant ones. In the online stage, we employ a progressive step-level reinforcement learning strategy, where truncated chains guide the model to infer missing steps and refine its reasoning, thereby sharpening its focus on critical steps and overcoming the limitations of static supervision. Extensive experiments on multimodal reasoning benchmarks show that our method achieves state-of-the-art performance while delivering superior training and inference efficiency. Code is available at https://github.com/kk0013/CritiCuRL.