arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ActReview:基于反驳引导的训练数据与评分标准奖励用于可操作的同行评审生成

ActReview: Rebuttal-Guided Training Data and Rubric Rewards for Actionable Peer Review Generation

Yiling Ma, Yilun Zhao, Sihong Wu, Ziyu Chen, Manasi Patwardhan, Arman Cohan

arXiv 2609.09076首次发表:更新:

发表机构

Yale University; University of Chicago; TCS Research(耶鲁大学; 芝加哥大学; 塔塔咨询服务研究中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对大模型预投稿自我审查中反馈缺乏可操作性的问题,提出ActReview框架,利用反驳引导构建训练数据并设计评分奖励,生成诊断与修改建议,优于现有模型。

AI 中文摘要

随着大型语言模型越来越多地用于投稿前的自我审查,人们对反馈的需求日益增长,这种反馈不仅要指出弱点,还要引导作者进行具体的修改。我们将此研究为可操作的同行评审生成,并将其分解为两个子任务:诊断性声明生成和修改建议生成。我们引入了ActReview,一个基于反驳引导的后训练框架,将针对论文的具体诊断与具体且有理有据的修改计划联系起来。我们的核心见解是,作者的反驳揭示了解决评审人关切的可操作行动,因此可以为面向修改的反馈提供潜在监督。从OpenReview上的真实评审-反驳线程中,我们通过将评审人的弱点与作者的回应对齐,并将由此产生的反馈基于局部论文证据,构建了ActReview-40K。我们使用多任务监督微调对Qwen3-8B-Base进行后训练,随后使用基于候选感知、弱点特定评分标准的GRPO。我们还引入了ActReview-Bench,一个包含1,000个实例的人工策划基准,用于评估诊断质量和修改有用性。实验表明,ActReview在可操作性和基于证据方面优于先前专门的评审生成模型,同时与基于强提示的LLM保持竞争力。人工评估证实了修改有用性的提升,同时揭示了在技术准确性方面仍存在的差距,进一步的分析支持对未见论文的泛化性和跨独立评审者的稳健性。

英文摘要

As LLMs are increasingly used for pre-submission self-review, there is growing demand for feedback that not only identifies weaknesses but also guides authors toward concrete revisions. We study this as Actionable Peer-review Generation and decompose it into two subtasks: diagnostic claim generation and revision suggestion generation. We introduce ActReview, a rebuttal-guided post-training framework that connects paper-specific diagnoses to concrete, grounded revision plans. Our central insight is that author rebuttals reveal plausible actions for addressing reviewer concerns and can therefore provide latent supervision for revision-oriented feedback. From real review-rebuttal threads on OpenReview, we construct ActReview-40K by aligning reviewer weaknesses with author responses and grounding the resulting feedback in localized paper evidence. We post-train Qwen3-8B-Base with multi-task supervised fine-tuning followed by GRPO using candidate-aware, weakness-specific rubric rewards. We also introduce ActReview-Bench, a human-curated benchmark of 1,000 instances for evaluating diagnostic quality and revision usefulness. Experiments show that ActReview outperforms prior specialized review-generation models on actionability and grounding while remaining competitive with strong prompt-based LLMs. Human evaluation confirms improved revision usefulness while revealing a remaining gap in technical accuracy, and additional analyses support generalization to held-out papers and robustness across independent judges.

Comments50 pages, 20 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑