arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.37889cs.CVcs.LG

ReCAP:面向多模态持续指令微调的检索引导能力复用

ReCAP: Retrieval-Guided Capability Reuse for Multimodal Continual Instruction Tuning

Tao Hu, Zhinuo Zhou, Xialiang Tong, De-Chuan Zhan, Da-Wei Zhou

首次发表
浏览论文内容

中文总结 AI 辅助

ReCAP提出检索引导框架,利用外部知识构建知识库并复用能力模块,通过自适应子空间回收实现多模态持续指令微调中的稳定跨阶段能力复用,达到SOTA性能。

中文摘要 AI 辅助

多模态持续指令微调(MCIT)旨在使多模态大语言模型能够从顺序任务中获取新能力,同时保留先前学到的知识。现有方法主要通过约束参数更新或分离任务特定适配来缓解灾难性遗忘。然而,持续适配也可以受益于外部知识,这些知识提供领域特定信息和可复用的推理模式,以解决多样化的指令。例如,要回答“球体左侧有多少个红色立方体?”,领域知识可以提供关于物体和空间关系的相关概念,而推理知识可以指定有序操作,如物体识别、空间过滤和计数。尽管有这种潜力,如何在持续适配中利用外部知识在现有MCIT方法中仍未得到充分探索。为此,我们提出ReCAP,一个检索引导框架,利用外部知识在持续适配过程中引导能力复用。在每个持续阶段,ReCAP使用外部搜索和LLM,基于当前阶段的训练数据,逐步构建包含领域、推理和格式知识的知识库。对于每条指令,检索到的领域知识引导生成,而检索到的推理知识选择并排序能力模块,形成实例特定的能力路径。由于这些能力模块跨阶段复用,后续适配可以覆盖先前学习的参数。为了实现稳定的跨阶段复用,ReCAP引入自适应子空间回收,该机制用共享基和阶段特定核心参数化可复用能力模块,保护历史上重要的方向,同时回收剩余容量。在MCIT基准上的大量实验表明,ReCAP达到了最先进的性能。

英文摘要

Multimodal continual instruction tuning (MCIT) aims to enable multimodal large language models to acquire new capabilities from sequential tasks while preserving previously learned knowledge. Existing methods primarily mitigate catastrophic forgetting by constraining parameter updates or separating task-specific adaptations. However, continual adaptation can also benefit from external knowledge that provides domain-specific information and reusable reasoning patterns for solving diverse instructions. For example, to answer "How many red cubes are to the left of the sphere?", domain knowledge can provide relevant concepts about objects and spatial relations, while reasoning knowledge can specify ordered operations such as object recognition, spatial filtering, and counting. Despite this potential, how to leverage external knowledge for continual adaptation remains largely unexplored in existing MCIT methods. To this end, we propose ReCAP, a retrieval-guided framework that leverages external knowledge to guide capability reuse during continual adaptation. At each continual stage, ReCAP uses external search and an LLM to incrementally build a knowledge base of domain, reasoning, and format knowledge based on the current-stage training data. For each instruction, retrieved domain knowledge guides generation, while retrieved reasoning knowledge selects and orders capability modules to form an instance-specific capability path. As these capability modules are reused across stages, subsequent adaptation can overwrite previously learned parameters. To enable stable cross-stage reuse, ReCAP introduces adaptive subspace recycling, which parameterizes reusable capability modules with shared bases and stage-specific cores, protects historically important directions while recycling residual capacity. Extensive experiments on MCIT benchmarks show that ReCAP achieves SOTA performance.

发表机构

  • Nanjing University(南京大学)
  • Huawei Technologies(华为技术有限公司)

机构由 AI 辅助整理,请以论文原文为准。

↑