arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

人工智能在协助物理、天体物理和宇宙学科学研究中的能力I:文献综述

AI's Capability in Assisting Scientific Research in Physics, Astrophysics, and Cosmology I: Literature Review

Anamaria Hell, Kateryna Vovk, Veena Krishnaraj, Jia Liu, Kosuke Aizawa, Adrian E. Bayer, Linda Blot, Jessica Cowell, Suyog Garg, Jonathan Grée, Ben Horowitz, Masaya Ichikawa, Kanyuni Iemoto, Keigo Kondo, Zacharie Lorsin, Kevin McCarthy, Jamie Robinson, Miguel Ruiz-Granda, Leander Thiele, Ievgen Vovk, Mingshen Zhou

arXiv 2607.25672首次发表:更新:

AI 中文总结

研究大语言模型在物理等领域协助文献综述的能力,通过对照研究发现其与人类所选文献重叠率低,生成的参考文献可靠性待提升,2026年的ChatGPT Pro 5.5性能有显著改善。

AI 中文摘要

我们研究了大语言模型(LLMs)在协助科学研究文献综述方面的表现。对物理、天体物理和宇宙学领域的八个专家设计的研究项目进行了对照研究。人类专家和人工智能提示器并行执行相同的文献综述任务。比较人类与2025年年中LLMs(ChatGPT - 4o、ChatGPT Deep Research和Gemini)所选的相关文献,发现重叠率小(<6%)。评估了人工智能生成的候选参考文献的可靠性和完整性,区分了两种幻觉类型。发现2025年年中模型需系统验证,而2026年的ChatGPT Pro 5.5性能显著提升,单项目测试无幻觉。

英文摘要

We investigate how well large language models (LLMs) can assist with literature reviews for scientific research. We perform a controlled study of eight expert-conceived research projects across the areas of physics, astrophysics, and cosmology. Each project has a defined background and goal, and human experts and AI prompters are asked to perform identical literature review tasks in parallel. We compare the relevant literature selected by humans with that selected by mid-2025 LLMs (ChatGPT-4o, ChatGPT Deep Research, and Gemini). We find the overlap between human- and AI-selected references to be small ($<$6\%), indicating that AI models do not yet reproduce a competent expert search on their own, though they have the potential to complement literature searches by humans. We then assess the reliability and completeness of AI-generated candidate references, distinguishing two types of hallucination: fabrications (references to nonexistent papers) and metadata mismatches (real papers with one or more incorrect fields). We find that while fabricated references make up 3\% of the AI-generated references, 64\% are real papers with at least one incorrect field (title, author, year, journal, DOI, or link), indicating that the mid-2025 models require systematic verification. However, the performance is significantly improved for the 2026 model ChatGPT Pro 5.5, with a single-project test showing zero fabrication or metadata mismatches.

Comments12 pages, 3 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑