发表机构
Rutgers University; Ernest Mario School of Pharmacy; East Brunswick High School; Princeton Medical Center; New Jersey Medical School; Rutgers Cancer Institute(罗格斯大学; 欧内斯特·马里奥药学院; 东布伦瑞克高中; 普林斯顿医疗中心; 新泽西医学院; 罗格斯癌症研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究比较了六种LLM平台在环境科学文献检索中的误差,发现基于摘要的提示比全文提示准确性更高,且检索准确性受平台、期刊、提示类型和输出位置影响,总体呈中等且不一致水平。
AI 中文摘要
大型语言模型(LLMs)越来越多地被用于文献检索和综合。然而,尚不清楚它们在环境科学中是否能检索到准确的文献信息。因此,我们定量比较了广泛使用的LLM平台在检索与来自五本领先环境科学期刊(《能源与环境科学》、《自然可持续性》、《自然气候变化》、《柳叶刀行星健康》和《环境科学与技术》)于2024年至2025年发表的原始文章相关的参考文献时的误差。Claude、ChatGPT、Grok、DeepSeek、Perplexity和Gemini被用作LLM平台。LLMs为50篇随机选择的原始文章中的每一篇检索了10篇参考文献,使用文章的摘要或其全文作为提示。检索到的参考文献经过一个多指标评分比率评估,该比率结合了文献数据的有效性、Google Scholar链接、数字对象标识符、Scopus电子标识符和相关性评分(被引用或为索引文章),以及在所有指标上均失败的完全捏造的比例。仅摘要提示产生的准确性显著高于全文提示。在调整了期刊、平台和输出顺序后,这一优势在多水平混合效应多变量回归中得到了确认。来源期刊和参考文献在输出列表中的位置也与检索准确性独立相关,列表中位置靠后的参考文献与较低的准确性相关。这些发现表明,环境科学中LLM辅助的文献检索仍然具有中等准确性且总体上不一致,因平台、期刊、提示类型和输出位置的不同而有显著差异。基于摘要的提示,作为任务对齐的信息压缩,在文献检索中可能优于全文提示。在推广我们的发现时应谨慎。
英文摘要
Large language models (LLMs) are increasingly used for literature search and synthesis. However, it is unclear whether they retrieve accurate bibliographic information in environmental science. Therefore, we quantitatively compared the errors of widely used LLM platforms in retrieving references related to original articles from five leading environmental science journals (Energy and Environmental Science, Nature Sustainability, Nature Climate Change, Lancet Planetary Health, and Environmental Science and Technology) published in 2024 to 2025. Claude, ChatGPT, Grok, DeepSeek, Perplexity, and Gemini were used as the LLM platforms. LLMs retrieved 10 references for each of the 50 randomly selected original article using either the article's abstract or its full-text as prompt. The retrieved references were subject to a multimetric score ratio combining validity of bibliographic data, Google Scholar link, digital object identifier, Scopus Electronic Identifier and relevance score (cited by or being the index paper), and the proportion of complete fabrication that failed all metrics. Abstract-only prompt yielded significantly higher accuracy than full-text one. This advantage was confirmed in multilevel mixed-effect multivariable regression after adjusting for journal, platform, and output order. Source journal and the position of a reference within the output list were also independently associated with retrieval accuracy, with lower-listed references associated with lower accuracy. These findings suggest that LLM assisted literature retrieval in environmental science remains moderately accurate and overall inconsistent, varying significantly by platform, journal, prompt type, and output position. Abstract-based prompting, as task-aligned information compression, may outperform full-text one in literature retrieval. Caution should be used when generalizing our findings.