arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.13742cs.CV

Hyper-LLaVA:用于多模态持续指令调优的双曲不确定性感知模态平衡路由

Hyper-LLaVA: Hyperbolic Uncertainty-aware Modality-Balanced Routing for Multimodal Continual Instruction Tuning

Kunlun Xu, Yanqin Zhang, Wenwen Qiang, Jiahuan Zhou

首次发表
浏览论文内容

中文总结 AI 辅助

针对多模态持续指令调优中参数路由的模态内信息利用不足和模态间可靠性差异问题,提出双曲不确定性感知模态平衡路由(Hyper-LLaVA),通过双曲空间分布相似性和不确定性量化实现自适应模态平衡,大幅超越现有方法。

中文摘要 AI 辅助

多模态持续指令调优(MCIT)旨在利用增量累积的知识来处理多样任务的多模态输入,其中参数路由发挥着重要作用。最先进的方法依赖于样本到任务中心的相似性以及路由过程中等权重的跨模态融合。然而,此类解决方案面临两个根本缺陷:(1)在每个模态内,样本到任务中心的距离对于路由而言是次优的,因为丰富的任务内多样性信息未被充分利用。(2)不同模态在不同任务中表现出不同的可靠性,其中具有任务间模糊性的模态容易误导路由结果。为解决这些问题,我们提出了双曲不确定性感知模态平衡路由(Hyper-LLaVA),基于跨模态任务特征不确定性建模来提升参数路由能力。具体而言,为改进模态内任务匹配,Hyper-LLaVA在双曲空间中评估样本到任务的分布相似性。此外,为缓解不可靠模态带来的性能退化,Hyper-LLaVA量化每个模态内的任务匹配模糊性,以实现跨模态任务匹配的自适应平衡。基于互补的模态内和模态间任务匹配增强,我们的Hyper-LLaVA大幅优于最先进的方法。我们的源代码可在以下网址获取:此https URL

英文摘要

Multimodal Continual Instruction Tuning (MCIT) aims to exploit the incrementally accumulated knowledge to process multimodal inputs of diverse tasks, where parameter routing plays an important role. State-of-the-art methods rely on sample-to-task center similarity and cross-modal fusion with equal weight during routing. However, such solutions face two fundamental flaws: (1) Within each modality, the sample-to-task center distance is sub-optimal for routing since the abundant intra-task diversity information is underleveraged. (2) Different modalities exhibit varying reliability across tasks, where the modality with inter-task ambiguity can easily misguide the routing result. To address these problems, we propose Hyperbolic Uncertainty-aware Modality-Balanced Routing (Hyper-LLaVA) to improve parameter routing capacity based on cross-modality task feature uncertainty modeling. Specifically, to improve intra-modality task matching, Hyper-LLaVA accesses the sample-to-task distribution similarity in the Hyperbolic space. Besides, to alleviate the degradation brought by unreliable modalities, Hyper-LLaVA quantifies the task matching ambiguity within each modality to achieve adaptive balancing between task matching across modalities. Based on the complementary intra- and inter-modality task matching enhancement, our Hyper-LLaVA outperforms state-of-the-art approaches by large margins. Our source code is available at https://github.com/zhoujiahuan1991/ICML2026-Hyper-LLaVA

发表机构

  • Wangxuan Institute of Computer Technology, Peking University(北京大学王选计算机技术研究所)
  • Institute of Software Chinese Academy of Sciences(中国科学院软件研究所)
  • University of Chinese Academy of Sciences(中国科学院大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑