arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AppliedScientist:通过迭代AI评审实现自动化科学修订

AppliedScientist: Automated Scientific Revision Through Iterative AI Reviewing

Vidushee Vats, Karun Sharma, Shengzhi Li, Shichao Pei

arXiv 2609.14738首次发表:更新:

发表机构

University of Massachusetts Boston(马萨诸塞大学波士顿分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出闭环系统AppliedScientist,将AI科学家与独立AI评审结合,迭代修订被拒论文。实验表明评审引导修订优于固定提示自我修订,能解决85.3%执行弱点,但仅解决11.1%想法弱点,揭示迭代修订的局限。

AI 中文摘要

自动评审系统越来越常根据其生成的评审质量来评估。然而,一份评审只有在依据其行动能带来论文可衡量的改进时才有用。我们提出了AppliedScientist,一个闭环系统,将自主AI科学家与AI评审员相结合,并通过迭代修订来自多个研究子领域的被拒论文来评估该系统。为了模拟人类作者基于早期草稿进行构建的方式,AI科学家在修订过程中可以访问其先前版本。然而,为避免先前判断带来的偏见,每次评审都是独立生成的,评审员对先前的反馈或评分没有记忆。我们比较了三种修订设置:一种以原始会议评审初始化,一种以AI生成的评审初始化,以及使用每轮相同固定提示的自主自我修订。由于评审员既指导又评估修订,我们还使用Stanford Reviewer作为独立评估者来评估人工初始化的修订。评审员引导的修订始终比固定提示的自我修订改进更多,且Stanford Reviewer也对后续修订给出更高分数。AppliedScientist解决了150个执行相关弱点中的128个(85.3%),但仅解决了18个想法相关弱点中的2个(11.1%),这表明迭代修订在改进实验和实现方面有效,但很少改变对新颖性或重要性的担忧。

英文摘要

Automated reviewing systems are increasingly evaluated based on the quality of the reviews they produce. Yet a review is only useful if acting on it leads to a measurable improvement in the paper. We present AppliedScientist, a closed-loop system that couples an autonomous AI scientist with an AI reviewer, and evaluate it by iteratively revising rejected papers from a range of research subfields. To mirror how human authors build on earlier drafts, the AI scientist has access to its previous versions during revision. To avoid bias from prior judgments, however, each review is generated independently, with the reviewer having no memory of earlier feedback or scores. We compare three revision settings: one initialized with the original venue reviews, one initialized with AI-generated reviews, and autonomous self-revision using the same fixed prompt in every round. Because the reviewer both guides and evaluates the revision, we also assess the human-initialized revisions using Stanford Reviewer as an independent evaluator. Reviewer-guided revision consistently improves more than fixed-prompt self-revision, and Stanford Reviewer also assigns higher scores to later revisions. AppliedScientist resolves 128 of 150 execution-related weaknesses (85.3%), but only 2 of 18 idea-related weaknesses (11.1%), suggesting that iterative revision is effective at improving experiments and implementation, but rarely changes concerns about novelty or significance.

Commentsv2: Updated Table 3

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑