发表机构
Sodhana(索达纳公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文探究领域特定三元组微调能否将预训练嵌入模型适配为身份敏感检索模型,通过构建合成数据集评估后发现,该方法可显著提升模型区分真实匹配与高度相似非匹配的能力,为相关应用提供实用改进方案。
AI 中文摘要
通用文本嵌入模型旨在捕捉语义相似度,但未针对区分代表同一真实世界实体(如企业或个人)的实体记录进行优化。这一局限性会影响实体解析、重复记录检索等应用,此类场景中细微的文本差异可能保留或改变实体身份。本文研究领域特定的三元组微调能否将预训练嵌入模型适配为对身份敏感的检索模型。研究人员构建了包含身份保留变体与具有挑战性非匹配示例的企业及个人记录合成数据集,采用基于间隔的相似度评估方法,在微调前后对两种广泛使用的嵌入模型进行评估。结果显示,模型在区分真实匹配与高度相似非匹配方面取得显著提升,表明领域特定的三元组训练可有效重塑通用嵌入空间以用于实体检索。这些发现表明,针对性微调为数据质量管理与信息检索应用中改进嵌入模型提供了实用方法。
英文摘要
General-purpose text embedding models are designed to capture semantic similarity but are not optimised for distinguishing entity records that represent the same real-world business or person. This limitation affects applications such as entity resolution and duplicate record retrieval, where small textual differences may either preserve or change identity. This paper investigates whether domain-specific triplet fine-tuning can adapt pretrained embedding models for identity-sensitive retrieval. A synthetic dataset of business and person records was created with identity-preserving variations and challenging non-matching examples. Two widely used embedding models were evaluated before and after fine-tuning using a margin-based similarity evaluation. The results show substantial improvements in separating true matches from highly similar non-matches, demonstrating that domain-specific triplet training can effectively reshape general-purpose embedding spaces for entity retrieval. These findings suggest that targeted fine-tuning provides a practical approach for improving embedding models in data quality management and information retrieval applications.