arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.10525cs.CVcs.AI

动态上下文适配器:将历史信息高效注入视觉-语言模型

Dynamic Context Adapters: Efficiently Infusing History into Vision-and-Language Models

Yuhang Song, Bor-Jiun Lin, Jiaxu Liu, Te-Chuan Chiu, Anh Nguyen, Chun-Yi Lee

首次发表
浏览论文内容

中文总结 AI 辅助

针对视觉-语言模型整合历史上下文时的计算与信息损失问题,提出动态上下文适配器(DCA),实现高效历史注入,降低注意力FLOPs超25%、内存13%,提升长时序任务性能。

中文摘要 AI 辅助

历史上下文整合是视觉-语言模型(VLMs)在序列决策任务中面临的核心挑战。当前VLMs独立处理视觉输入,对需要时序理解的下游应用造成关键局限;直接将历史帧融入Transformer输入会产生二次注意力复杂度与过高内存消耗,现有方法存在计算膨胀或通过时序压缩造成大量信息损失的显著缺陷。为解决这些挑战,我们提出动态上下文适配器(Dynamic Context Adapter,DCA),一种用于预训练VLMs的新型上下文注入方法。该方法采用固定大小、动态压缩的内存来保留历史语义,无需帧拼接,可衔接静态VLMs与循环策略,在预训练模型中实现内存能力的同时保持计算效率。DCA使注意力浮点运算量(FLOPs)降低超过25%,内存节省13%,并提升了长时序任务的性能。

英文摘要

Historical context integration presents a fundamental challenge for Vision-Language Models (VLMs) in sequential decision-making tasks. Current VLMs process visual inputs independently, which creates critical limitations for downstream applications that require temporal understanding. Direct incorporation of historical frames into Transformer inputs produces quadratic attention complexity and excessive memory consumption. Existing approaches suffer from significant drawbacks: computational inflation or substantial information loss through temporal compression. To address these challenges, we introduce Dynamic Context Adapter (DCA), a novel context injection approach for pretrained VLMs. Our method employs fixed-size, dynamically compressed memory to preserve historical semantics without frame concatenation. DCA bridges static VLMs and recurrent policies and enables memory capabilities in pretrained models while maintaining computational efficiency. DCA achieves over $25\%$ reduction in attention FLOPs and $13\%$ memory savings while improving performance on long-horizon tasks.

发表机构

  • University of Liverpool(利物浦大学)
  • National Tsinghua University(国立清华大学)
  • Imperial College London(伦敦帝国学院)
  • National Taiwan University(国立台湾大学)

机构由 AI 辅助整理,请以论文原文为准。

↑