InternReviewer与InternAdvocate:面向同行评审与反驳的智能体强化学习的客观奖励与评估
InternReviewer & InternAdvocate: Objective Reward and Evaluation for Agentic Reinforcement Learning in Peer Review and Rebuttal
浏览论文内容
中文总结 AI 辅助
本研究提出InternReviewer与InternAdvocate框架,构建学术数据集与arXiv检索工具,采用多维度标准的客观奖励系统优化智能体RL,训练后智能体推理深度与引用准确性显著提升。
中文摘要 AI 辅助
生成同行评审、反驳等专业学术内容,需要领域推理与事实依据的复杂协同。本研究提出用于开发和评估专用学术智能体的综合框架InternReviewer与InternAdvocate。我们首先构建大规模高质量学术数据集,并集成高效arXiv检索工具以支持主动证据收集。为优化这些智能体,我们实现由统一客观指标与奖励系统驱动的智能体强化学习(RL)范式,该系统采用多维度标准(包括参考锚定语义对齐、结构合规性)及严格验证机制(核对引用与实时交互日志以消除幻觉),避免基于主观模型评判的偏差。实验结果显示,在该闭环框架中训练的智能体,其推理深度与引用准确性均显著提升。
英文摘要
Generating professional scholarly content, such as peer reviews and rebuttals, requires an intricate synergy between domain reasoning and factual grounding. This work presents a comprehensive framework for the development and evaluation of specialized scholarly agents, InternReviewer and InternAdvocate. We first establish a large-scale, high-quality scholarly dataset and integrate a high-efficiency arXiv retrieval tool to enable active evidence gathering. To optimize these agents, we implement an agentic Reinforcement Learning (RL) paradigm driven by a unified objective metric and reward system. This system avoids the biases of subjective model-based judging by employing multi-dimensional criteria, including reference-anchored semantic alignment, structural compliance, and a strict verification mechanism that cross-checks citations against real-time interaction logs to eliminate hallucinations. Experimental results demonstrate that agents trained within this closed-loop framework exhibit significant improvements in reasoning depth and citation accuracy.
发表机构
- Shanghai AI Laboratory(上海人工智能实验室)
- Beijing Jiaotong University(北京交通大学)
- Tsinghua University(清华大学)
- Renmin University of China(中国人民大学)
机构由 AI 辅助整理,请以论文原文为准。