arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.12060cs.CV

回顾过去,展望未来:按需视觉记忆用于高效多模态推理

Look Back, Think Ahead: Visual Memory on Demand for Efficient Multimodal Reasoning

Yicheng Xue, Han Wu, Jufeng Yang, Minjing Dong, Xinghao Chen, Hanting Chen, Jianyuan Guo

首次发表
浏览论文内容

中文总结 AI 辅助

针对多模态大语言模型处理高分辨率图像长视觉标记序列时固定压缩上下文无法适配推理中视觉证据需求变化的问题,提出轻量框架ViMoD,通过DART和TRACE实现按需视觉记忆,在Qwen3-VL-4B上以20%视觉标记预算提升推理性能且仅需少量额外参数。

中文摘要 AI 辅助

处理高分辨率图像的长视觉标记序列会使多模态大语言模型(MLLM)的多步推理计算成本高昂。现有的一次性剪枝和聚合方法会在解码前将视觉标记压缩为固定上下文,但随着推理推进,视觉证据需求会发生变化,固定压缩上下文难以保留各阶段所需的全部细节。为应对这一挑战,我们提出ViMoD这一轻量框架,在推理需求演变时,既能维持紧凑的视觉上下文,又能保留对原始细粒度证据的访问权限。区域标记的可变形聚合(DART)学习内容自适应分组和聚合能力,构建与可恢复原始细(Fine)标记关联的紧凑粗(Coarse)表示;自适应上下文证据的时序路由(TRACE)整合解码历史,预测后续证据需求,选择、保留或替换活跃的细标记组。选定的细标记会增强冻结主干中的持久粗上下文,实现特定阶段的证据访问,无需持续关注所有视觉标记。在Qwen3-VL-4B上,ViMoD在20%的目标视觉标记预算下,于全部8个推理基准上优于所有评估的基线,比最强的一次性基线的平均归一化分数提升39.0%,且仅需冻结主干中0.0546%的额外可训练参数。

英文摘要

Processing long visual token sequences from high-resolution images makes multi-step reasoning computationally expensive for multimodal Large Language Models (MLLMs). Existing one-shot pruning and aggregation methods compress visual tokens into a fixed context before decoding. However, visual evidence needs can shift as reasoning unfolds, making it difficult for a fixed compressed context to retain all the details needed across stages. To address this challenge, we propose ViMoD, a lightweight framework that maintains a compact visual context while preserving access to original fine-grained evidence as reasoning needs evolve. Deformable Aggregation of Region-wise Tokens (DART) learns content-adaptive groups and aggregation capacities, constructing compact Coarse representations linked to recoverable original Fine tokens. Temporal Routing for Adaptive Contextual Evidence (TRACE) integrates decoding history to anticipate upcoming evidence needs and select, retain, or replace active Fine-token groups. Selected Fine tokens augment the persistent Coarse context in the frozen backbone, enabling stage-specific evidence access without continuously attending to all visual tokens. On Qwen3-VL-4B, ViMoD outperforms all evaluated baselines on all eight reasoning benchmarks at a 20% target visual token budget, improving the mean normalized score by 39.0% over the strongest evaluated one-shot baseline. These gains are achieved with only 0.0546% additional trainable parameters relative to the frozen backbone.

发表机构

  • City University of Hong Kong(香港城市大学)
  • Zhejiang University(浙江大学)
  • Peking University(北京大学)
  • Nankai University(南开大学)
  • Huawei Technologies(华为技术有限公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑