arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

OPLD:用于多模态推理的在线策略隐层蒸馏

OPLD: On-Policy Latent Distillation for Multimodal Reasoning

Shoutai Zhu, Tianyang Xu, Bin Sun, Mingyuan Xu, Yu Liu, Qinzhen Guo

arXiv 2607.28154首次发表:更新:

AI 中文总结

本文提出OPLD框架,将多模态CoT的推理能力迁移至隐层表征,经多模态基准实验,其性能优于现有隐层推理方法,达到当前最优。

AI 中文摘要

交错式多模态思维链(CoT)通过将辅助视觉证据融入中间推理过程,可提升视觉推理能力,但现有方法仍受限于外部定义的推理轨迹与视觉操作,难以发展灵活且抽象的视觉思维。近期,基于隐层的推理为将中间计算内化为连续表征提供了有前景的方向,然而现有视觉-隐层方法主要通过与压缩的辅助视觉特征对齐来监督隐层状态,将其视为视觉观测的代理而非主动推理状态,导致模型仅能捕获提供的证据,无法完全内化多模态CoT引发的抽象推理过程。本文提出OPLD(在线策略隐层蒸馏),这一简单框架可将特权多模态CoT引发的推理能力迁移至隐层推理表征。在多种多模态基准上开展的大量实验表明,OPLD始终优于现有隐层推理方法,并在多个基准上达到了当前最优性能。结果表明,在推理过程层面监督隐层表征,相比传统的特征级对齐,为多模态隐层推理提供了更有效的范式。

英文摘要

Interleaved multimodal Chain-of-Thought (CoT) improves visual reasoning by incorporating auxiliary visual evidence into intermediate reasoning. However, existing approaches remain constrained by externally defined reasoning traces and visual operations, limiting their ability to develop flexible and abstract visual thinking. Reasoning with latent has recently offered a promising direction by internalizing intermediate computation into continuous representations. Nevertheless, existing visual-latent methods mainly supervise latent states through alignment with compressed auxiliary visual features, treating them as proxies for visual observations rather than active reasoning states. Consequently, they capture the provided evidence but fail to fully internalize the abstract reasoning process induced by multimodal CoT. In this paper, we propose OPLD (On-Policy Latent Distillation), a simple framework that transfers the reasoning capability induced by privileged multimodal CoT into latent reasoning representations. Extensive experiments on diverse multimodal benchmarks demonstrate that OPLD consistently outperforms existing latent reasoning methods and achieves state-of-the-art performance on multiple benchmarks. The results suggest that supervising latent representations at the reasoning-process level provides a more effective paradigm for multimodal latent reasoning than conventional feature-level alignment.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑