发表机构
The Hong Kong University of Science and Technology; Alibaba Group; Huazhong University of Science and Technology; Nanjing University(香港科技大学; 阿里巴巴集团; 华中科技大学; 南京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对多模态嵌入中余弦对比限制语义表达、点积训练不稳定的问题,提出两阶段渐进框架PACE,逐步扩展表示与参数空间,并引入焦点嵌入损失,实验验证其有效性。
AI 中文摘要
多模态嵌入模型将异构输入编码到共享嵌入空间,从而能够跨模态和任务进行高效相似度计算。大多数现有方法优化基于余弦的对比目标,这促进了稳定的训练,但将语义兼容性限制在角度几何上,使得嵌入范数无法作为额外的语义信号。然而,直接优化更具表达力的点积相似度(同时利用角度和范数信息)在性能上不如基于余弦的训练,并表现出不稳定的训练动态。我们将这种差异归因于过早的优化空间扩展,表现为表示空间中的角度-范数纠缠和方向各向异性,并且全参数微调进一步加剧了这一问题。在本文中,我们提出了PACE,一个两阶段框架,逐步扩展表示空间和可训练参数空间。第一阶段结合基于余弦的目标与低秩适应,在受限优化空间中建立可靠的角度几何。第二阶段切换到点积相似度和全参数微调,使嵌入方向和范数能够共同编码语义信息。我们进一步引入了焦点嵌入损失,一种置信度自适应目标,它降低具有高正检索置信度的查询的权重,同时强调具有竞争性负样本的模糊查询。跨多个骨干规模和多样化的多模态嵌入任务的实验一致验证了PACE的有效性。
英文摘要
Multimodal embedding models encode heterogeneous inputs into a shared embedding space, enabling efficient similarity computation across modalities and tasks. Most existing methods optimize cosine-based contrastive objectives, which promote stable training but restrict semantic compatibility to angular geometry, precluding embedding norms from serving as an additional semantic signal. However, directly optimizing the more expressive dot-product similarity, which leverages both angular and norm information, underperforms cosine-based training and exhibits unstable training dynamics. We attribute this discrepancy to premature optimization-space expansion, manifested as angular--norm entanglement and directional anisotropy in the representation space and further compounded by full-parameter fine-tuning. In this paper, we propose PACE, a two-stage framework that progressively expands both the representation and trainable parameter spaces. Stage I combines cosine-based objective with low-rank adaptation to establish a reliable angular geometry within constrained optimization spaces. Stage II switches to dot-product similarity and full-parameter fine-tuning, enabling embedding directions and norms to jointly encode semantic information. We further introduce Focal Embedding Loss, a confidence-adaptive objective that downweights queries with high positive retrieval confidence while emphasizing ambiguous queries with competitive negatives. Experiments across multiple backbone scales and diverse multimodal embedding tasks consistently validate the effectiveness of PACE.