发表机构
Shanghai AI Laboratory; University of California, Los Angeles; The Hong Kong University of Science and Technology(上海人工智能实验室; 加州大学洛杉矶分校; 香港科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出AutoSupervision,利用《自然·通讯》5.6万篇文章的评审记录构建数据集,发现LLM刻画评审关切表现较好但证据验证是瓶颈,为AI辅助科学工作流提供反馈循环闭合方案。
AI 中文摘要
大型语言模型(LLM)的最新进展使AI系统能够辅助科学研究与同行评审,但可靠的AI辅助科学工作流中,验证评审意见是否能带来有意义、有证据支撑的手稿改进这一关键能力仍未得到充分探索。我们提出AutoSupervision,它通过基于事实的证据评估科学手稿修订是否真正解决了评审者的关切。AutoSupervision将透明的同行评审记录作为天然监督源,其中评审意见明确科学关切、作者回复描述声称的解决方案、修订后的手稿提供修改证据。给定评审意见、作者回复与修订后的手稿,模型需刻画评审者的关切、判定关切是否已解决并识别手稿中的支持证据。我们从56000篇《自然·通讯》文章及对应评审记录构建AutoSupervision,随后在LLM上开展实验、 ablation研究与案例研究。结果显示,LLM在刻画评审者关切方面表现良好,其中GPT-5.5取得0.754的分数,但基于证据的验证仍是主要瓶颈,表现最佳的模型仅达到0.501。
英文摘要
Recent advances in large language models (LLMs) have enabled AI systems to assist scientific research and peer review. However, an essential capability for reliable AI-assisted scientific workflows remains underexplored: verifying whether reviewer feedback leads to meaningful and evidence-supported manuscript improvements. We introduce AutoSupervision, which evaluates whether scientific manuscript revisions genuinely address reviewer concerns through grounded evidence. AutoSupervision leverages transparent peer-review records as a natural source of supervision, where reviewer comments specify scientific concerns, author responses describe claimed resolutions, and revised manuscripts provide evidence of changes. Given reviewer comments, author responses, and revised manuscripts, models must characterize reviewer concerns, determine whether concerns have been addressed, and identify supporting manuscript evidence. We construct AutoSupervision from 56,000 Nature Communications articles and corresponding review records. Then we conducted experiments on LLMs, the ablation study, and the case study. Our results show that while LLMs perform well in characterizing reviewer concerns, with GPT-5.5 achieving a score of 0.754, evidence-based verification remains the primary bottleneck, with the best-performing model reaching only 0.501.