arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

幻觉检测器:一种用于在AtomGPT.org上检测科学文献中幻觉的混合大语言模型和语义学者工具调用

Hallucination Detector: A hybrid LLM and Semantic Scholar tool calling for detecting hallucination in scientific literature on AtomGPT.org

Harichandana Neralla, Jaehyung Lee, Aldo H. Romero, Kamal Choudhary

arXiv 2607.09774首次发表:更新:

AI 中文总结

研究针对大语言模型用于科学写作时出现的虚构参考文献问题,提出结合大语言模型字段提取与语义学者结构化检索的AtomGPT参考文献检查器,经基准测试能可靠标记多数幻觉引用,保障文献引用可信度。

AI 中文摘要

大语言模型如今常用于科学写作,这带来了一种更隐蔽的失败:虚构参考文献。伪造作者、虚假数字对象标识符(DOI)、错误分配的标识符以及合并多篇真实文章元素的引用,正大量插入稿件中,传统同行评审难以处理。近期审核发现此类参考文献已通过评审流程进入已发表文献。因此,能以现代内容生产速度和规模运行的自动验证成为必要保障。本文介绍并评估了AtomGPT参考文献检查器,它通过结合大语言模型字段提取和语义学者的结构化检索来验证文献引用,针对每条引用进行字段提取、检索匹配真实论文并评分,判断其可信度。通过与一组来自已接受的NeurIPS 2025论文的经外部策划的确认幻觉引用进行基准测试,发现该工具能可靠地标记出绝大多数此类引用。

英文摘要

Large language models are now commonly used as partners in scientific writing, and this shift has brought a subtler type of failure: made-up references. Fabricated authors, bogus DOIs, wrongly assigned identifiers, and citations that merge elements from multiple genuine articles are now being inserted into manuscripts at a volume that traditional peer review was never meant to handle. Recent audits reveal that such references have already slipped through the review process and made their way into the published literature, including leading journals and conferences. Automated verification that operates at the speed and scale of modern content production has therefore become a necessary safeguard rather than a convenience. This work presents and evaluates the AtomGPT reference checker (https://atomgpt.org/hallucination_detector), an open, web-accessible tool that verifies citations against the scholarly literature by combining large-language-model field extraction with structured retrieval from Semantic Scholar. For each reference, the tool extracts the bibliographic fields, retrieves the closest matching real papers, and scores the agreement across title, authorship, and venue to produce a graded judgment of whether a citation is trustworthy, partially supported, or likely fabricated. We benchmark the tool against an externally curated set of confirmed hallucinated citations from accepted NeurIPS 2025 papers and find that it reliably flags the great majority of them.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑