发表机构
Max Planck Institute for Intelligent Systems; ELLIS Institute Tübingen(马克斯·普朗克智能系统研究所; 图宾根ELLIS研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究大语言模型引用验证器的失败模式,通过HALLMARK基准测试评估多种验证器,发现误报率决定验证器是否可部署,并具体指出三种失败模式,强调误报率是部署瓶颈,未检测到的伪造对科学记录代价更高。
AI 中文摘要
大语言模型如今常用于撰写文献综述和辅助学术写作,这使得虚假引用风险增加,如GPTZero在NeurIPS 2025的录用论文中发现53篇有幻觉引用的论文。基于规则和大语言模型的验证器不断涌现,但缺乏共享基准来比较它们并提供详细的失败诊断。我们用HALLMARK(幻觉基准)填补这一空白,它包含2526个BibTeX条目,涵盖14种幻觉类型、三个难度等级、每个条目六个诊断子测试以及一个抗污染的留出划分。我们在其上评估了DOI查找基线、前沿大语言模型的零样本、工具增强代理以及我们自己基于规则的协同设计验证器bibtex - updater。整个基准测试的一个一致结果是:误报率而非召回率决定验证器是否可部署。HALLMARK通过三种失败模式使其具体化:代理查找提高召回率但增加误报;在现实的场地基础率下,误报率的数量级差异而非召回率决定验证器的标记大多是真正的发现还是大多是噪声;大多数大语言模型对在其训练截止日期之后发表的论文过度标记,只有两个最新截止日期的模型将误报率保持在分布水平附近(我们将此信号作为描述性报告,因为它与这些条目的可能召回率混淆)。因此,误报率是部署瓶颈,但未检测到的伪造对科学记录来说仍然是代价更高的错误。
英文摘要
Large language models (LLMs) now routinely draft literature reviews and assist with academic writing, which means a higher risk of fabricated references: GPTZero found 53 papers with hallucinated citations among NeurIPS 2025's accepted set. Rule- and LLM-based verifiers are emerging, but no shared benchmark compares them and gives detailed failure diagnostics. We close that gap with HALLMARK (Hallucination benchmark): 2,526 BibTeX entries spanning 14 hallucination types, three difficulty tiers, six diagnostic sub-tests per entry, and a contamination-resistant held-out split. On it we evaluate a DOI-lookup baseline, frontier LLMs zero-shot, tool-augmented agents, and our own rule-based, co-designed verifier bibtex-updater. Across the benchmark one result is consistent: the false-positive rate, not recall, decides whether a verifier is deployable. HALLMARK makes it concrete through three failure modes: agentic lookups buy recall but inflate false positives; at a venue-realistic base rate, the order-of-magnitude spread in false-positive rates (FPRs) -- not recall -- governs whether a verifier's flags are mostly true catches or mostly noise; and most LLMs over-flag papers published past their training cutoff, where only the two latest-cutoff models hold their false-positive rate near in-distribution levels (a signal we report as descriptive, since it is confounded with possible recall of those entries). Thus FPR is the deployment bottleneck, but an undetected fabrication remains the costlier error for the scientific record.
Comments59 pages, 7 figures, 40 tables. Benchmark and code: https://github.com/rpatrik96/hallmark (v1.2.0); verification tool: https://github.com/rpatrik96/bibtexupdater (v1.5.0)