arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.25881cs.CLastro-ph.COastro-ph.IMcs.HCgr-qc

人工智能在协助物理、天体物理和宇宙学科学研究中的能力II:项目规划与提案评估

AI's Capability in Assisting Scientific Research in Physics, Astrophysics, and Cosmology II: Project Planning and Proposal Evaluation

Jia Liu, Veena Krishnaraj, Kateryna Vovk, Kosuke Aizawa, Adrian E. Bayer, Linda Blot, Jessica Cowell, Suyog Garg, Jonathan Grée, Anamaria Hell, Ben Horowitz, Ma… 展开作者

Jia Liu, Veena Krishnaraj, Kateryna Vovk, Kosuke Aizawa, Adrian E. Bayer, Linda Blot, Jessica Cowell, Suyog Garg, Jonathan Grée, Anamaria Hell, Ben Horowitz, Masaya Ichikawa, Kanyuni Iemoto, Keigo Kondo, Zacharie Lorsin, Kevin McCarthy, Jamie Robinson, Miguel Ruiz-Granda, Leander Thiele, Ievgen Vovk, Mingshen Zhou

首次发表
浏览论文内容

中文总结 AI 辅助

研究大语言模型在科学项目规划与提案评估中的作用,人类和三个当代模型生成计划,再由人类与前沿模型盲评,结果显示当前模型能产出类似人工的计划,但人工智能评审员更青睐人工智能生成的提案,部署时需谨慎。

中文摘要 AI 辅助

我们研究了大语言模型(LLMs)在协助科学项目规划和提案评估方面的效果。人类研究人员和三个当代大语言模型(ChatGPT、Claude和DeepSeek;2025年年中版本,使用默认工具访问)为八个物理、天体物理和宇宙学领域专家构思的研究项目独立生成单页项目计划。32份提案由四名人类评审员和两个前沿大语言模型(Claude Opus 4.8和ChatGPT Pro 5.5)使用四方面评估标准进行盲评。评审员还需判断提案作者是人还是人工智能。总体而言,人类评审员对人工和人工智能撰写的提案评分相近,而两个人工智能评审员对人工智能撰写的提案评分比人工撰写的高约一分(五分制)。人类评审员分别有72%和79%的时间正确识别出人工和人工智能撰写的提案,而两个人工智能评审员均100%正确分类所有32份提案。结果表明,当前大语言模型能生成在人类评审员眼中与人工撰写相当的项目计划,但人工智能评审员对人工智能生成的提案有系统偏好,并建议在提案准备和评估中广泛部署大语言模型时要谨慎。

英文摘要

We investigate how well large language models (LLMs) can assist scientific project planning and proposal evaluation. One-page project plans were independently generated for eight expert-conceived research projects in physics, astrophysics, and cosmology by human researchers and three contemporary LLMs (ChatGPT, Claude, and DeepSeek; mid-2025 models, used with their default tool access). The resulting 32 proposals were blindly evaluated by four human reviewers and two newer frontier LLMs (Claude Opus 4.8 and ChatGPT Pro 5.5) using a four-aspect evaluation rubric. Reviewers were also asked to identify whether each proposal was written by a human or an AI. Human reviewers rated human- and AI-written proposals similarly overall, whereas both AI reviewers scored AI-written proposals about one point higher (on a five-point scale) than human-written proposals. Human reviewers correctly identified human- and AI-written proposals 72% and 79% of the time, respectively, while both AI reviewers correctly classified all 32 proposals (100%). These results suggest that current LLMs can produce project plans comparable to human-written ones in the eyes of human reviewers, but that AI reviewers show a systematic preference for AI-generated proposals. Our results suggest caution when deploying LLMs widely in proposal preparation and evaluation.

补充信息

↑