AI 中文总结
研究大语言模型在物理等领域协助文献综述的能力,通过对照研究发现其与人类所选文献重叠率低,生成的参考文献可靠性待提升,2026年的ChatGPT Pro 5.5性能有显著改善。
AI 中文摘要
我们研究了大语言模型(LLMs)在协助科学研究文献综述方面的表现。对物理、天体物理和宇宙学领域的八个专家设计的研究项目进行了对照研究。人类专家和人工智能提示器并行执行相同的文献综述任务。比较人类与2025年年中LLMs(ChatGPT - 4o、ChatGPT Deep Research和Gemini)所选的相关文献,发现重叠率小(<6%)。评估了人工智能生成的候选参考文献的可靠性和完整性,区分了两种幻觉类型。发现2025年年中模型需系统验证,而2026年的ChatGPT Pro 5.5性能显著提升,单项目测试无幻觉。
英文摘要
We investigate how well large language models (LLMs) can assist with literature reviews for scientific research. We perform a controlled study of eight expert-conceived research projects across the areas of physics, astrophysics, and cosmology. Each project has a defined background and goal, and human experts and AI prompters are asked to perform identical literature review tasks in parallel. We compare the relevant literature selected by humans with that selected by mid-2025 LLMs (ChatGPT-4o, ChatGPT Deep Research, and Gemini). We find the overlap between human- and AI-selected references to be small ($<$6\%), indicating that AI models do not yet reproduce a competent expert search on their own, though they have the potential to complement literature searches by humans. We then assess the reliability and completeness of AI-generated candidate references, distinguishing two types of hallucination: fabrications (references to nonexistent papers) and metadata mismatches (real papers with one or more incorrect fields). We find that while fabricated references make up 3\% of the AI-generated references, 64\% are real papers with at least one incorrect field (title, author, year, journal, DOI, or link), indicating that the mid-2025 models require systematic verification. However, the performance is significantly improved for the 2026 model ChatGPT Pro 5.5, with a single-project test showing zero fabrication or metadata mismatches.
Comments12 pages, 3 figures