arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

多模态测试时自适应的表示编辑

Representation Editing for Multimodal Test-Time Adaptation

Longfei Huang, Xiangyu Wu, Yang Yang

arXiv 2609.32263首次发表:更新:

AI 中文总结

针对多模态测试时自适应中中间表示错位问题,提出FIRE方法,通过逐层表示编辑与频域混合优化,实现跨模态对齐,提升预测可靠性。

AI 中文摘要

多模态测试时自适应(TTA)旨在利用未标记的测试数据,在线调整预训练的多模态模型以适应跨模态的分布偏移,在现实应用中展现出广泛潜力。然而,现有方法主要侧重于调整融合特征以弥合源域与目标域之间的差距,缺乏对中间表示错位的显式控制,而中间表示错位正是分布偏移下性能下降的关键驱动因素。在本工作中,我们从表示工程的角度应对这一挑战。与以往就地更新融合权重的TTA方法不同,我们提出了四元数傅里叶表示编辑器(FIRE),一种新颖的多模态TTA方法,直接编辑语义丰富的中间表示。具体而言,我们首先将表示编辑器引入单模态编码器的每个中间层,实现单模态表示的逐层校准。为了进一步增强低秩编辑子空间的多样性和稳定性,每个表示编辑器通过快速傅里叶变换进行频域混合,以构建结构化基。此外,我们引入多级自适应目标来优化这些编辑器,共同促进跨模态语义对齐、源-目标统计对齐以及非对称预测一致性。通过这种方式,FIRE为融合产生对齐的单模态表示,并进一步提高预测可靠性。在两个广泛使用的多模态基准上,针对多种损坏类型的广泛实验证明了FIRE相对于现有TTA方法的优越性。

英文摘要

Multimodal test-time adaptation (TTA) aims to adapt a pretrained multimodal model online to distribution shift across modalities using unlabeled test data, showing broad potential in real-world applications. However, existing methods primarily focus on adjusting fused features to bridge the source-target gap, lacking explicit control over intermediate representation misalignment, which is a key driver of performance drop under distribution shift. In this work, we tackle this challenge from the perspective of representation engineering. Unlike previous TTA methods that update fusion weights in place, we propose FourIer Representation Editor (FIRE), a novel multimodal TTA approach that directly edits semantically rich intermediate representations. Specifically, we first adopt representation editors into each intermediate layer of the unimodal encoders, enabling layer-wise calibration of unimodal representations. To further enhance the diversity and stability of the low-rank editing subspaces, each representation editor performs frequency domain mixing via the fast Fourier transform to construct structured bases. Moreover, we introduce multi-level adaptation objectives to optimize these editors, jointly promoting cross-modal semantic alignment, source-target statistical alignment, and asymmetric prediction consistency. In this way, FIRE yields aligned unimodal representations for fusion and further improves prediction reliability. Extensive experiments on two widely used multimodal benchmarks under various corruption types demonstrate the superiority of FIRE over existing multimodal TTA methods.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑