当AI评审训练AI评审者:科学判断崩溃及其缓解
When AI Reviews Train AI Reviewers: Scientific-Judgment Collapse and Mitigation
浏览论文内容
中文总结 AI 辅助
本文研究AI评审递归训练导致科学判断崩溃的问题,提出TrustReviewer系统,通过训练时精选语料和测试时激活引导缓解该问题,保持判断多样性。
中文摘要 AI 辅助
大型语言模型(LLMs)越来越多地参与科学评估,既作为自动评审者,也作为人类评审者的助手。随着模型生成的评审进入公共数据和未来训练语料库,AI同行评审可能变得递归:后来的评审者从早期模型产生的判断中学习。我们在受控环境中研究了这一反馈循环的一个步骤。从Llama 3.1 8B开始,我们首先在2018--2023年的官方ICLR评审上微调一个评审者,然后在2024年ICLR数据上训练四个后继模型,这些数据包含系统变化的官方评审和模型生成评审的混合。我们的研究表明,引入合成评审压缩了评分分布,并减少了同论文和语料库级别的语义多样性。我们将这种模式称为\u201c科学判断崩溃\u201d。为了缓解这一失败模式,我们引入了TrustReviewer,一个开源的基于LLM的系统,用于生成AI和机器学习论文的同行评审。TrustReviewer在两个互补阶段进行干预。对于训练时预防,我们在一个精选语料库上单阶段训练核心评审者,该语料库旨在减少低质量和语义退化的监督。对于测试时纠正,配对激活引导旨在进一步缓解崩溃判断的残余倾向,无需进一步训练或额外专家注释。总之,这些结果刻画了递归评审者训练的具体风险,并为在AI辅助科学评估中保持判断多样性和改进推荐对齐提供了实际干预措施。
英文摘要
Large language models (LLMs) increasingly participate in scientific evaluation, both as automated reviewers and as assistants to human reviewers. As model-generated reviews enter public data and future training corpora, AI peer review can become recursive: later reviewers learn from judgments produced by earlier models. We study one step of this feedback loop in a controlled setting. Starting from Llama 3.1 8B, we first fine-tune a reviewer on official ICLR reviews from 2018--2023 and then train four successor models on ICLR 2024 data with systematically varied mixtures of official and model-generated reviews. Our study shows that introducing synthetic reviews compresses rating distributions and reduces both same-paper and corpus-level semantic diversity. We call this pattern $\textbf{scientific-judgment collapse}$. To mitigate this failure mode, we introduce $\textbf{TrustReviewer}$, an open-source LLM-based system for generating peer reviews of AI and machine learning papers. TrustReviewer intervenes at two complementary stages. For training-time prevention, we train the core reviewer in a single stage on a curated corpus designed to reduce low-quality and semantically degenerate supervision. For test-time correction, paired activation steering aims to further mitigate residual tendencies toward collapsed judgments without further training or additional expert annotation. Together, these results characterize a concrete risk of recursive reviewer training and provide practical interventions for preserving judgment diversity and improving recommendation alignment in AI-assisted scientific evaluation.
发表机构
- University of Maryland, College Park(马里兰大学学院公园分校)
机构由 AI 辅助整理,请以论文原文为准。