arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.05980cs.LG

简单变换在文本嵌入模型间的迁移性研究

How Far Do Simple Transformations Translate Across Text Embedding Models?

Sid Ali Hamideche, Louis-Adrien Dufrène, Quentin Lampin, Guillaume Larue

首次发表
浏览论文内容

中文总结 AI 辅助

该研究探究简单变换在异构文本嵌入模型间的迁移性,通过9种不同模型的多维度评估发现,异构嵌入空间并非普遍可通过简单映射关联。

中文摘要 AI 辅助

我们研究简单变换能否在异构文本嵌入模型间迁移表示。理解独立训练的模型如何组织语义信息,可实现AI间的潜在通信而无需解码为人类可读文本。聚焦线性映射等轻量转换器,我们在超越简化基准的真实文本场景中,测试文献中关于潜在普遍性的假设。针对9种架构、池化策略、训练目标不同的嵌入模型,我们使用CKA、下游迁移、保真度和检索评估兼容性。简单转换器可恢复有意义的共享结构,支持部分兼容对的迁移,但在其他对中表现极差。兼容性同时取决于架构、训练目标、池化和数据分布。总体而言,结果表明异构嵌入空间并不像部分文献所暗示的那样,普遍可通过简单映射关联。

英文摘要

We investigate whether simple transformations can translate representations across heterogeneous text embedding models. Understanding how independently trained models organize semantic information is an enabler for AI-to-AI latent communication without decoding into human-readable text. Focusing on lightweight translators such as linear mappings, we test the literature hypothesis of latent universality in a realistic text setting beyond simplified benchmarks. Across nine embedding models differing in architecture, pooling strategy, and training objective, we evaluate compatibility using CKA, downstream transfer, fidelity, and retrieval. Simple translators recover meaningful shared structure and support transfer for some compatible pairs, but fail sharply for others. Compatibility depends jointly on architecture, training objective, pooling, and data distribution. Overall, the results show that heterogeneous embedding spaces are not universally related by simple mappings as often suggested in some literature.

↑