arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AlphaWiSE:用于连续多模态表示学习的自适应权重插值

AlphaWiSE: Adaptive Weight Interpolation for Continual Multimodal Representation Learning

Sarthak Jain, Qiran Hu, Zhen Zhu, Yaoyao Liu

arXiv 2607.15094首次发表:更新:

发表机构

Google DeepMind(谷歌DeepMind)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究连续多模态表示学习中跨模态对齐被破坏的问题,提出AlphaWiSE方法,通过事后权重空间插值,由两个冻结源检查点拟合系数实现插值检查点,在多模态检索实验中相比基线有持续改进。

AI 中文摘要

多模态模型如CLIP为跨模态检索学习共享嵌入空间,但持续适应顺序到达的数据会破坏早期阶段获得的跨模态对齐。传统连续学习方法返回单个检查点,使每个检索方向都面临相同的稳定性-可塑性权衡。我们提出了AlphaWiSE,一种事后权重空间插值方法,它由两个冻结的源检查点组成。对于由其检查点键标识的每个对齐参数张量,AlphaWiSE拟合一个由所有张量条目共享的标量插值系数。这些系数在较小的示例内存上拟合,并用于实现一个插值检查点。部署的模型具有与任何一个源检查点相同的架构和参数数量,不需要额外的推理时间。在音频-图像-文本检索上的大量实验表明,在多个检索方向和评估指标上,相较于强大的连续学习基线有持续的改进。

英文摘要

Multimodal models such as CLIP learn a shared embedding space for cross-modal retrieval, but continual adaptation to sequentially arriving data can disrupt the cross-modal alignment acquired from earlier phases. Conventional continual-learning methods return a single checkpoint, which commits every retrieval direction to the same stability-plasticity trade-off. We propose AlphaWiSE, a post-hoc weight-space interpolation method that composes two frozen source checkpoints. For each aligned parameter tensor identified by its checkpoint key, AlphaWiSE fits one scalar interpolation coefficient shared by all tensor entries. The coefficients are fitted on a smaller exemplar memory and used to materialize one interpolated checkpoint. The deployed model has the same architecture and parameter count as either source checkpoint, which does not require additional inference time. Extensive experiments on audio-image-text retrieval show consistent improvements over strong continual-learning baselines across multiple retrieval directions and evaluation metrics.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑