arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.24209cs.MM

面向统一框架内多模态视听学习任务的任务解耦低秩适配方法

Task-disentangled Low-Rank Adaptation for Versatile Audio-visual Multi-modal Learning Tasks within a Unified Framework

Hanyu Xuan, Mengqi Zhang, Junjun Mao, Fei Wang, Kun Li, Guanghui Yue, Zhiliang Wu, Hehe Fan

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出统一框架下的任务解耦低秩适配(LoRA)机制,解决多视听任务联合训练的干扰问题,在多项视听任务上优于现有统一模型及部分任务特定模型。

中文摘要 AI 辅助

受人类多模态感知的启发,视听多模态学习(AVMML)整合听觉与视觉信息以利用互补的跨模态线索,实现更鲁棒、全面的场景感知。现有研究大多孤立处理每项AVMML任务,这与人类处理多样感知的统一认知能力形成鲜明对比。然而,对多项AVMML任务进行简单联合训练常因复杂的任务间关系而出现相互干扰。为解决该问题,本文提出一种可同时适配多样AVMML任务的统一框架。具体而言,利用大语言模型强大的表征与泛化能力,设计任务解耦低秩适配(LoRA)机制,实现任务特定知识与任务共享知识的动态整合,进而促进有效的多任务协作。所提任务解耦LoRA包含三个组件:任务通用低秩矩阵、任务特定调制矩阵与跨任务协作专家,分别捕获通用视听知识、解耦任务特定模式、挖掘固有任务间关联。通过从模型与任务视角统一显式协作,本文方法在多项AVMML任务上不仅优于现有统一视听模型,且在部分AVMML任务上的表现优于大多数任务特定模型。

英文摘要

Inspired by human multi-modal perception, Audio-Visual Multi-Modal Learning (AVMML) integrates auditory and visual information to leverage complementary cross-modal cues, enabling more robust and comprehensive scene perception. Existing studies predominantly tackle each AVMML task in isolation, which stands in stark contrast to humans' unified cognitive capacity for handling versatile perception. However, naive joint training across multiple AVMML tasks often suffers from mutual interference, arising from the intricate inter-task relationships. To address this, we propose a unified framework that simultaneously accommodates versatile AVMML tasks. Specifically, benefiting from powerful representation and generalization capabilities of large language models, we design a task-disentangled Low-Rank Adaptation (LoRA) mechanism that enables dynamic integration of both task-specific and task-shared knowledge, thereby facilitating effective multi-task collaboration. The proposed task-disentangled LoRA comprises three components: a task-general low-rank matrix, task-specific modulation matrices, and cross-task collaboration experts, which respectively capture universal audio-visual knowledge, decouple task-specific pattern, and exploit inherent inter-task correlations. By unifying explicit collaboration from both model and task perspectives, our approach not only surpasses existing unified audio-visual models across multiple AVMML tasks, but also outperforms most task-specific models on certain AVMML tasks.

↑