arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.28612cs.AI

InternReviewer与InternAdvocate:面向同行评审与反驳的智能体强化学习的客观奖励与评估

InternReviewer & InternAdvocate: Objective Reward and Evaluation for Agentic Reinforcement Learning in Peer Review and Rebuttal

Xuerui Su, Liya Guo, Qizhi Pei, Qipeng Guo, Zhongbo Tian, Lijun Wu, Kai Chen, Zun Wang

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出InternReviewer与InternAdvocate框架,构建学术数据集与arXiv检索工具,采用多维度标准的客观奖励系统优化智能体RL,训练后智能体推理深度与引用准确性显著提升。

中文摘要 AI 辅助

生成同行评审、反驳等专业学术内容,需要领域推理与事实依据的复杂协同。本研究提出用于开发和评估专用学术智能体的综合框架InternReviewer与InternAdvocate。我们首先构建大规模高质量学术数据集,并集成高效arXiv检索工具以支持主动证据收集。为优化这些智能体,我们实现由统一客观指标与奖励系统驱动的智能体强化学习(RL)范式,该系统采用多维度标准(包括参考锚定语义对齐、结构合规性)及严格验证机制(核对引用与实时交互日志以消除幻觉),避免基于主观模型评判的偏差。实验结果显示,在该闭环框架中训练的智能体,其推理深度与引用准确性均显著提升。

英文摘要

Generating professional scholarly content, such as peer reviews and rebuttals, requires an intricate synergy between domain reasoning and factual grounding. This work presents a comprehensive framework for the development and evaluation of specialized scholarly agents, InternReviewer and InternAdvocate. We first establish a large-scale, high-quality scholarly dataset and integrate a high-efficiency arXiv retrieval tool to enable active evidence gathering. To optimize these agents, we implement an agentic Reinforcement Learning (RL) paradigm driven by a unified objective metric and reward system. This system avoids the biases of subjective model-based judging by employing multi-dimensional criteria, including reference-anchored semantic alignment, structural compliance, and a strict verification mechanism that cross-checks citations against real-time interaction logs to eliminate hallucinations. Experimental results demonstrate that agents trained within this closed-loop framework exhibit significant improvements in reasoning depth and citation accuracy.

发表机构

  • Shanghai AI Laboratory(上海人工智能实验室)
  • Beijing Jiaotong University(北京交通大学)
  • Tsinghua University(清华大学)
  • Renmin University of China(中国人民大学)

机构由 AI 辅助整理,请以论文原文为准。

↑