AI 中文总结
该研究针对多模态持续指令调优中专家路由的任务识别问题,构建含34个任务的FLEX基准,提出MCIL框架,适配CIL方法后使LoRA匹配和MacroScore分别提升最高16.3个百分点、4.6个点
AI 中文摘要
多模态持续指令调优(Multimodal Continual Instruction Tuning, MCIT)使多模态大语言模型能够按顺序获取新任务,同时保留先前学习的能力。近期许多方法维护任务特定的LoRA专家,并在推理时将每个输入路由至一个或多个专家,但专家路由背后的任务识别问题仍未得到充分探索。我们表明,在广泛使用的MCIT基准上,路由几乎已饱和;泄露任务身份的文本指纹,以及仅含4-10个任务、竞争专家较少的短序列,共同掩盖了长程路由问题。为揭示这一挑战,我们引入FLEX(Fingerprint-reduced Long-horizon Expert eXamination),这是一个包含34个任务的长程MCIT基准,其文本指纹已被弱化。FLEX将具有相似指令和答案格式但视觉与知识领域不同的任务分组,标准化它们的外部模板,并在大得多的专家池上评估路由。关键是,我们将渐进式LoRA路由表述为软任务作为类的多模态类增量学习(Multimodal Class-Incremental Learning, MCIL):每个任务定义一个增量路由类,其完整分数分布提供LoRA混合权重,硬路由是其离散特例。FLEX揭示了这一日益严峻的任务识别挑战,而MCIL表述为将CIL方法迁移至专家路由提供了原则性接口。我们实例化PureLoRA作为受控基线,并将四种CIL方法适配至四个MCIT框架,且不修改其LoRA专家或生成流水线。我们的插件路由器将严格LoRA匹配提升了多达16.3个百分点,整体MacroScore提升了多达4.6个点。代码可在以下URL获取:this https URL
英文摘要
Multimodal Continual Instruction Tuning (MCIT) enables multimodal large language models to acquire new tasks sequentially while retaining previously learned capabilities. Many recent methods maintain task-specific LoRA experts and route each input to one or more experts at inference. Yet the task-identification problem underlying expert routing remains under-explored. We show that routing is nearly saturated on widely used MCIT benchmarks. Textual fingerprints that leak task identity and short 4--10-task sequences with few competing experts jointly obscure the long-horizon routing problem. To expose this challenge, we introduce FLEX (Fingerprint-reduced Long-horizon Expert eXamination), a 34-task long-horizon MCIT benchmark with weakened textual fingerprints. FLEX groups tasks with similar instruction and answer formats but diverse visual and knowledge domains, normalizes their outer templates, and evaluates routing over a substantially larger expert pool. Crucially, we formulate progressive-LoRA routing as soft task-as-class Multimodal Class-Incremental Learning (MCIL): each task defines an incremental routing class, whose complete score distribution supplies the LoRA mixture weights, with hard routing as a discrete special case. FLEX exposes this expanding task-identification challenge, while the MCIL formulation provides a principled interface for transferring CIL methods to expert routing. We instantiate PureLoRA as a controlled baseline and adapt four CIL methods to four MCIT frameworks without modifying their LoRA experts or generation pipelines. Our plug-in routers improve strict LoRA matching by up to 16.3 percentage points and overall MacroScore by up to 4.6 points. Code is available at: https://github.com/RINC-CL/FLEX