TRAM:利用轨迹衍生辅助记忆增强多模态推理
TRAM: Enhancing Multimodal Reasoning with Trajectory-Derived Auxiliary Memory
浏览论文内容
中文总结 AI 辅助
TRAM是一种无需训练的多模态推理增强方法,通过轨迹衍生辅助记忆整合推理信息,在4种MLRM变体的8个基准测试中提升了多模态推理性能。
中文摘要 AI 辅助
多模态大推理模型(Multimodal Large Reasoning Models,MLRMs)在需要视觉理解和多步推理的任务上已取得出色性能。然而,随着推理轨迹变长,模型对上下文早期建立的信息的利用能力可能下降,增加推理错误的风险。现有方法主要通过在整个推理过程中维持视觉接地来解决该问题,但推理还会将视觉观察转化为任务特定的关系、约束和中间结论,其影响可能会在长轨迹中减弱。归因分析表明,推理正确性并非仅由图像属性一致区分,而是与轨迹是否在各阶段保留并整合此类推理衍生信息更紧密相关。受此启发,我们提出TRAM(TRajectory-derived Auxiliary Memory),这是一种无需训练的方法,通过模型自身推理轨迹衍生的辅助记忆通路增强标准解码过程。TRAM将已完成的推理整合为紧凑的潜在记忆,通过快慢循环流在线更新,并通过轻量残差通路将其反馈至选定的解码器层。在8个基准测试上对4种MLRM变体开展的实验显示,TRAM在无需额外训练的情况下,提升了标准解码在数学、科学及通用视觉推理任务上的性能。
英文摘要
Multimodal Large Reasoning Models (MLRMs) have achieved strong performance on tasks requiring visual understanding and multi-step inference. However, as reasoning trajectories grow, models may become less effective at using information established earlier in the context, increasing the risk of reasoning errors. Existing approaches primarily address this problem by sustaining visual grounding throughout reasoning. However, reasoning also transforms visual observations into task-specific relations, constraints, and intermediate conclusions whose influence may weaken over long trajectories. Our attribution analysis suggests that correctness is not consistently separated by image attribution alone, but is more closely associated with whether trajectories retain and integrate such reasoning-derived information across stages. Motivated by this, we introduce TRAM (TRajectory-derived Auxiliary Memory), a training-free method that augments standard decoding with an auxiliary memory pathway derived from the model's own reasoning trajectory. TRAM consolidates completed reasoning into a compact latent memory, updates it online through fast and slow recurrent streams, and feeds it back into selected decoder layers through a lightweight residual pathway. Experiments across four MLRM variants on eight benchmarks show that TRAM improves performance over vanilla decoding on mathematical, scientific, and general visual reasoning tasks without additional training.