通过合成查询探测映射嵌入模型间的相似度空间
Mapping Similarity Spaces across Embedding Models with Synthetic Query Probing
- Pegasystems(pegasystems公司)
- Leiden University(莱顿大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究针对嵌入模型间相似度分数不可直接比较的问题,提出合成查询探测方法,通过学习分数分布映射实现跨模型相似度空间对齐,提升阈值可移植性,为嵌入可比性分析提供可扩展框架。
AI中文摘要:
检索增强生成(RAG)系统依赖相似度分数来检索相关内容,但由于不同嵌入模型具有不同的几何属性,分数无法直接比较,这使得模型迁移变得复杂,并限制了阈值的复用。我们研究如何通过学习分数分布之间的映射(而非嵌入之间的映射)来关联相似度分数。我们引入合成查询探测(Synthetic Query Probing),即从文档生成查询以创建受控的查询-文本块对,从而实现对跨模型相似度行为的大规模、无参考分析。我们在多种嵌入配置上评估该方法,并使用线性、保序和分位数映射学习分数转换函数。在SciFact和一个专有语料库上的实验表明,尽管模型在排序上大体一致,但其绝对分数存在系统性偏差。学习到的映射部分对齐了这些空间并提高了阈值的可移植性,其中保序回归(isotonic regression)表现最佳。我们的结果强调了跨模型校准的必要性,并将合成查询探测定位为分析嵌入可比性的可扩展框架。
英文摘要:
Retrieval-Augmented Generation systems rely on similarity scores to retrieve relevant content, yet scores are not directly comparable across embedding models due to differing geometric properties, complicating model migration and limiting threshold reuse. We study how similarity scores can be related by learning mappings between score distributions rather than embeddings. We introduce Synthetic Query Probing, generating queries from documents to create controlled query-chunk pairs, enabling large-scale, reference-free analysis of cross-model similarity behavior. We evaluate the approach on multiple embedding configurations and learn score conversion functions using linear, isotonic, and quantile mappings. Experiments on SciFact and a proprietary corpus show that while models largely agree on rankings, their absolute scores exhibit systematic distortions. Learned mappings partially align these spaces and improve threshold portability, with isotonic regression performing best. Our results highlight the need for cross-model calibration and position Synthetic Query Probing as a scalable framework for analyzing embedding comparability.