arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MuRA:用于高效且有效的测试时视觉-语言泛化的多秩适配

MuRA: Multi-Rank Adaptation for Efficient and Effective Test-Time Vision-Language Generalization

Gengyuan Liu, Nanzhou Wang, Chang Liu, Qinwen Wu, Zhenhao Wang, Jiacong Wang, Bokui Chen, Xiangyang Ji

arXiv 2608.03885首次发表:更新:

AI 中文总结

针对测试时视觉-语言模型因静态秩配置导致的性能瓶颈,提出多秩适配(MuRA)框架,动态选择适配模块,在多基准测试中达最优准确率且降低计算内存开销。

AI 中文摘要

视觉-语言模型展现出出色的零样本能力,但在分布偏移下会出现显著的性能下降。尽管通过低秩适配(Low-Rank Adaptation)进行的测试时适配(Test-Time Adaptation,TTA)提供了参数高效的解决方案,我们发现当前方法存在一个根本性瓶颈:依赖静态秩配置。由于视觉输入固有地具有不同的信息密度,固定秩会导致不可避免的优化折中,进而在复杂场景下出现欠拟合,在简单场景下出现过拟合。为弥合这一差距,我们提出多秩适配(Multi-Rank Adaptation,MuRA),这是一种基于 token 级视觉复杂度动态选择并融合不同容量适配模块的新型框架。MuRA 协同多秩正交分解(Multi-Rank Orthogonal Decomposition)以提供更优的、保留知识的初始化,并结合带连续路由器更新的统一组件融合(Unified Component Fusion with Continuous Router Updating)以持续学习语义到秩的映射。此外,我们提供严格的理论证明,从数学上证明这种自适应机制的必要性和梯度稳定性。关键的是,MuRA 的动态设计在最深层视觉层表现尤为突出,利用了最短的梯度反向传播路径。大量实验表明,MuRA 在广泛的域泛化和跨数据集基准测试中达到了最先进的准确率,同时显著降低了计算和内存开销。

英文摘要

Vision-language models exhibit remarkable zero-shot capabilities but suffer significant performance degradation under distribution shifts. While test-time adaptation (TTA) via Low-Rank Adaptation offers a parameter-efficient solution, we identify a fundamental bottleneck in current methods: the reliance on static rank configurations. Because visual inputs inherently possess varying information densities, a fixed rank forces an inevitable optimization compromise, leading to underfitting on complex scenes and overfitting on simple ones. To bridge this gap, we propose Multi-Rank Adaptation (MuRA), a novel framework that dynamically selects and fuses adaptation modules of varying capacities based on token-level visual complexity. MuRA synergizes Multi-Rank Orthogonal Decomposition to provide a superior, knowledge-preserving initialization, and Unified Component Fusion with Continuous Router Updating to sustainably learn semantic-to-rank mappings. Furthermore, we provide rigorous theoretical justifications mathematically proving the necessity and gradient stability of this adaptive mechanism. Crucially, MuRA's dynamic design uniquely thrives at the deepest visual layer, capitalizing on the shortest gradient backpropagation path. Extensive experiments demonstrate that MuRA achieves state-of-the-art accuracy across extensive domain generalization and cross-dataset benchmarks while significantly reducing both computational and memory overhead.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑