arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2410.18105cs.IRcs.AIcs.CL

利用实体关系图与模型感知对比采样提升文档检索的嵌入准确率

Improving Embedding Accuracy for Document Retrieval Using Entity Relationship Maps and Model-Aware Contrastive Sampling

  • Fifth Dimension AI(第五维度人工智能公司)

机构由 AI 辅助整理,请以论文原文为准。

Thea Aviss

更新

AI总结:

提出APEX-Embedding-7B,通过结构化实体关系图的收敛前中断微调与模型感知对比采样,提升长上下文文档RAG检索的事实聚焦能力并取得90.86% rank@1。

AI中文摘要:

本文提出 APEX-Embedding-7B(Advanced Processing for Epistemic eXtraction),这是一个拥有70亿参数的仅解码器文本特征提取模型,专门面向文档检索增强生成(RAG)任务设计。我们的方法采用两种训练技术,从而在事实聚焦方面产生涌现式改进:(1)使用结构化实体关系图作为训练数据输入的收敛前中断微调:旨在转移模型注意力,并使其偏向事实内容而非语义风格——尽管没有直接针对纯文本进行训练,这仍提升了纯文本性能;(2)模型感知对比采样,根据基础模型的能力,创建一个均衡且均匀分布的硬负样本与软负样本整理映射。这种组合方法带来显著改进,将纯文本查询/文档对检索提升至在我们的评估中实现90.86%的绝对rank@1准确率(相比排名第二的领先模型提高6.26%),并且与查询和文档文本的纯文本相比,训练数据输入上下文大小平均减少37.71%。基于评估,我们的模型在较长上下文文档检索任务的文本特征提取方面确立了新的最先进标准。

英文摘要:

In this paper we present APEX-Embedding-7B (Advanced Processing for Epistemic eXtraction), a 7-billion parameter decoder-only text Feature Extraction Model, specifically designed for Document Retrieval-Augmented Generation (RAG) tasks. Our approach employs two training techniques that yield an emergent improvement in factual focus: (1) Pre-convergence interrupted fine-tuning using Structured Entity Relationship Maps as training data input: designed to shift the model's attention and create a bias towards factual content rather than semantic style - this enhances plain text performance despite not being directly trained for it; and (2) Model-Aware Contrastive Sampling, creating a balanced and evenly distributed collation map of hard and soft negatives directly informed by the base model's competency. This combined methodology yields significant improvements, enhancing plain text query/document pair retrieval to achieve an absolute rank@1 accuracy of 90.86% (an increase of 6.26% compared to the next leading model) in our evaluation, and reducing training data input context size by an average of 37.71% compared to plain text for both queries and document texts. Based on our evaluations, our model establishes a new state-of-the-art standard in text feature extraction for longer context document retrieval tasks.

补充信息

↑