生物医学参考文献生成在26个大语言模型中仍不可靠
Biomedical Reference Generation Remains Unreliable across 26 Large Language Models
浏览论文内容
中文总结 AI 辅助
本研究评估26个大语言模型在生物医学参考文献生成中的可靠性,发现编造率高达55.4%,且无模型在全部字段上正确率超过54.6%,强调使用前必须验证。
中文摘要 AI 辅助
背景。大语言模型越来越多地被用于帮助撰写生物医学文本,但可能编造出不存在的参考文献。大语言模型这样做的频率尚未得到充分表征。方法。我们提示了来自八家开发者的26个语言模型(2023年至2026年),为十个领域的69个生物医学段落中的每一个提供缺失的参考文献。参考文献被分类为可验证(具有可解析标识符的真实论文)、部分匹配(没有可解析标识符的真实论文)、编造(没有匹配的索引论文)或拒绝(模型拒绝提供参考文献)。只有当参考文献可验证且其期刊、年份和列出的作者与所引用论文匹配时,该参考文献才被认为在每个评估的文献字段中都是正确的。结果。编造率从10.2%(Claude Opus 4.8,拒绝了52.1%的提示)到98.4%(Ministral 3B,未产生任何可验证的参考文献)不等。Claude Opus 4.6和Claude Sonnet 4.5产生了相似比例的可验证参考文献(分别为77.6%和76.6%),但在可评估作者的参考文献中,分别有78.7%和28.7%正确列出了作者,并且在每个评估字段中分别有54.6%和19.9%的响应完全正确。GPT-5.5在48.1%的响应中每个字段都正确。在所有模型中,55.4%的响应是编造的,14.9%在每个字段中都正确。在首次发布于2026年的五个测试模型中,相应的比例分别为35.3%和31.8%。结论。编造仍然普遍,没有模型在超过54.6%的响应中每个评估的文献字段都正确。能够识别真实论文的模型仍可能错误陈述其元数据,因此使用模型辅助生成的参考文献在使用前需要验证。
英文摘要
Background. Large language models are increasingly used to help write biomedical text but may fabricate references to nonexistent work. How often large language models do so is not well characterized. Methods. We prompted 26 language models from eight developers (2023 to 2026) to supply a missing reference for each of 69 biomedical passages across ten domains. References were classified as verifiable (real paper with a resolving identifier), partial matches (real paper without a resolving identifier), fabricated (no matching indexed paper), or declined (the model refused to supply a reference). A reference was considered correct in every evaluated bibliographic field only when it was verifiable and its journal, year, and listed authors matched those of the cited paper. Results. Fabrication ranged from 10.2% (Claude Opus 4.8, which declined 52.1% of prompts) to 98.4% (Ministral 3B, which produced no verifiable reference). Claude Opus 4.6 and Claude Sonnet 4.5 produced similar proportions of verifiable references (77.6% and 76.6%) but named authors correctly in 78.7% and 28.7% of author-evaluable verifiable references, respectively, and were correct in every evaluated field in 54.6% and 19.9% of responses. GPT-5.5 was correct in every field in 48.1%. Across all models, 55.4% of responses were fabricated and 14.9% were correct in every field. Among the five tested models first released in 2026, the corresponding proportions were 35.3% and 31.8%, respectively. Conclusions. Fabrication remained common, and no model was correct in every evaluated bibliographic field in more than 54.6% of responses. Models that identify real papers may still misstate their metadata, so references produced with model assistance require verification before use.
发表机构
- Columbia University(哥伦比亚大学)
- VNS Health(VNS健康机构)
- Tel Aviv Sourasky Medical Center(特拉维夫苏拉斯基医疗中心)
- University of Eastern Finland(东芬兰大学)
- Wellbeing Services County of North Savo(北萨沃福祉服务县)
- Wellbeing Services County of Southwest Finland(西南芬兰福祉服务县)
机构由 AI 辅助整理,请以论文原文为准。