发表机构
The Chinese University of Hong Kong; National University of Singapore; Yale University; The Hong Kong University of Science and Technology (Guangzhou)(香港中文大学; 新加坡国立大学; 耶鲁大学; 香港科技大学(广州))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对双曲多模态持续学习中的几何失真问题,提出HMCL方法,通过最近可行校正保留共享等距,在16任务流上提升性能并减少漂移。
AI 中文摘要
现有的持续学习方法保护参数、重放示例或欧几里得特征子空间。当应用于双曲多模态模型时,它们没有显式保留联合编码模态内相似性、跨模态对应关系和语义层次结构的洛伦兹几何;因此,顺序更新可以保留任务分数,同时仍然扭曲先前学习的关系。我们通过双曲多模态持续学习(HMCL)来解决这一差距。我们表明,保留旧的多模态几何等同于将所有模态限制在一个共享的双曲等距变换上,这产生了一族可行的第一阶参数变化。我们制定了一个联合最近可行(CA)校正,该校正保留与候选模态更新最佳匹配的共享旋转;其最小旋转(MR)特例将此旋转固定为零。两种变体都纠正了AdamW实现的位移,任务锚定限制了任务内累积,同时保留了学习自由度。在具有三个双曲骨干网络的统一16任务分类-检索流中,HMCL相对于顺序微调和四个持续学习基线提高了最终性能和向后迁移;HMCL-CA在每个骨干网络上给出了最高的总体得分。模态扩展流证实了检索增益。表示分析发现径向、角度、跨模态和配对距离漂移减少了81.2%至95.5%;ImageNet-WordNet结果显示了更好的语义祖先和径向层次结构。
英文摘要
Existing continual-learning methods protect parameters, replayed examples, or Euclidean feature subspaces. When applied to hyperbolic multimodal models, they do not explicitly preserve the Lorentz geometry that jointly encodes within-modality similarity, cross-modal correspondence, and semantic hierarchy; sequential updates can therefore retain task scores while still distorting previously learned relations. We address this gap with Hyperbolic Multimodal Continual Learning (HMCL). We show that preserving the old multimodal geometry amounts to restricting all modalities to one shared hyperbolic isometry, which induces a family of admissible first-order parameter changes. We formulate a joint closest-admissible (CA) correction that retains the shared rotation best matching the candidate modal updates; its minimal-rotation (MR) special case fixes this rotation to zero. Both variants correct the displacement realized by AdamW, and task anchoring bounds within-task accumulation while preserving learning freedom. Across a unified 16-task classification-retrieval stream with three hyperbolic backbones, HMCL improves final performance and backward transfer over sequential fine-tuning and four continual-learning baselines; HMCL-CA gives the highest Overall score on every backbone. A modality-extended stream confirms the retrieval gains. Representation analyses find 81.2 to 95.5 percent less radial, angular, cross-modal, and paired-distance drift; ImageNet-WordNet results show better semantic ancestry and radial hierarchy.
Comments49 pages, 10 figures, 11 tables