面向高能与天体粒子物理的检索增强生成
Towards Retrieval Augmented Generation in High-Energy and Astroparticle Physics
- Universite de Montreal(蒙特利尔大学)
- University of Maryland(马里兰大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对高能与天体粒子物理文献庞大难查的问题,开发开源RAG流水线,混合检索并重排序约23万篇论文,用Gemma-4-E4B生成带引用的摘要报告,超越传统关键词搜索。
AI中文摘要:
高能与天体粒子物理文献已发展成为一个庞大的知识体系。全面理解这些文献是研究中的重要步骤,但也是一项具有挑战性的任务。在取得有意义的进展之前,重要的是要确定已完成的工作并识别开放问题。然而,鉴于出版物的数量庞大,这一过程正变得越来越困难。针对狭窄主题进行聚焦式文献调研尤其具有挑战性。我们开发了一个开源的检索增强生成(RAG)流水线,以生成基于arXiv论文数据库的、带有引用的简洁且有针对性的摘要报告,使研究人员、编辑和审稿人能够快速熟悉该领域的任何角落。具体而言,我们对约23万篇hep-ph和astro-ph.HE论文进行嵌入,以实现同时捕捉关键词匹配和语义相似性的混合检索。融合后的候选集随后通过交叉编码器重排序器进行精炼。检索后,论文被传递给大语言模型(LLM),我们使用Google DeepMind的开源Gemma-4-E4B模型。该LLM首先充当法官评估相关性,然后综合出一份连贯且有依据的报告并附上引用。这种现代AI驱动的方法使我们能够超越arXiv或Inspire-HEP上经典关键词搜索的能力。
英文摘要:
The high-energy and astroparticle physics literature has grown into an enormous body of knowledge. Gaining a comprehensive understanding of this literature is an important step in research, but is a challenging task. Before making meaningful advances, it is important to establish what has already been done and identify open questions. However, this process is becoming increasingly difficult given the sheer volume of publications. Conducting a focused literature survey on a narrow topic is particularly challenging. We develop an open-source Retrieval Augmented Generation (RAG) pipeline to produce concise, targeted summary reports with citations, grounded in the database of arXiv papers, enabling researchers, editors, and referees to rapidly orient themselves within any corner of the field. Specifically, we embed approximately 230K hep-ph and astro-ph.HE papers to enable a hybrid retrieval that captures both keyword matches and semantic similarity. The fused candidate set is then refined using a cross-encoder reranker. Following retrieval, the papers are passed to a Large Language Model (LLM), for which we use Google DeepMind's open-source Gemma-4-E4B model. The LLM first acts as a judge to evaluate the relevance and then synthesizes a coherent and grounded report with references. This modern AI-driven approach allows us to go beyond the capabilities of classical keyword search on arXiv or Inspire-HEP.