arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

REPAIR:通过事实验证的迭代细化解决科学检索器中的长尾混淆问题

REPAIR: Resolving Long-Tail Confusion in Scientific Retrievers via Fact-Verified Iterative Refinement

Yerim Oh, Gunhee Kim

arXiv 2609.18262首次发表:更新:

发表机构

Seoul National University(首尔大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

REPAIR通过自进化数据增强框架,迭代诊断长尾概念并利用API证据扩展和难负样本挖掘,显著提升科学密集检索器在材料科学和生物医学基准上的性能。

AI 中文摘要

科学信息的精确检索从根本上受到科学语料库中长尾概念和高事实敏感性的限制。这些挑战常常限制密集检索器的有效性以及易产生幻觉的LLM增强方法。为了解决这一问题,我们提出了REPAIR,一个用于科学密集检索器的自进化数据增强框架。REPAIR通过循环执行长尾概念诊断、API引导的证据扩展以及通过难负样本挖掘进行区分,迭代地合成训练数据以解决知识缺口。这一过程有效地将检索锚定在事实现实中,以解决细粒度的区分。大量实验表明,REPAIR在九个材料科学和生物医学基准上显著优于19个强基线。我们的工作强调,针对长尾缺陷进行诊断和事实增强数据对于稳健的科学检索至关重要。

英文摘要

Precise retrieval of scientific information is fundamentally constrained by long-tailed concepts and high fact-sensitivity of scientific corpora. These challenges often limit the effectiveness of dense retrievers and hallucination-prone LLM augmentation. To address this, we present REPAIR, a self-evolving data augmentation framework for scientific dense retrievers. REPAIR iteratively synthesizes training data to address knowledge gaps by cycling through diagnosis of long-tail concepts, API-guided evidence expansion, and differentiation via hard negative mining. This process effectively grounds retrieval in factual reality to resolve fine-grained distinctions. Extensive experiments demonstrate that REPAIR significantly outperforms 19 strong baselines on nine materials science and biomedical benchmarks. Our work highlights that diagnosing and factually augmenting data to long-tail deficits is essential for robust scientific retrieval.

CommentsAccepted to EMNLP 2026 (Main Conference). 30 pages, 5 figures, 20 tables. Code: https://github.com/yerimoh/REPAIR

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑