arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PaperGym:以评分标准为中心的研究计划生成演化框架

PaperGym: Rubric-Centered Evolution for Research-Plan Generation

Yuhan Wang, Zhengxi Lu, Yuchen Yan, Kaitao Song, Wenqi Zhang, Weiming Lu, Jun Xiao, Yueting Zhuang, Yongliang Shen

arXiv 2608.31119首次发表:更新:

发表机构

Zhejiang University; Apple(浙江大学; 苹果公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出PaperGym框架,将论文转化为研究计划生成的训练环境,采用评分标准分阶段训练Qwen3系列模型,显著提升了研究规划能力,在多项基准上优于现有方法。

AI 中文摘要

研究规划是AI科学家的关键能力,但研究计划不存在可验证的答案,因此强化学习缺乏所需的环境——配对有评判者的任务。从科学论文中提取的评分标准可充当该评判者。然而现有流程存在缺陷:问题与评判标准来自同一内容,导致模型可通过改述获得奖励;且评分标准被压缩为每轮次的单一标量。本文提出PaperGym,一个将每篇研究论文转化为完整训练环境的统一框架。PaperGym利用论文结构:问题由研究目标与背景综合生成,评判标准则源自方法与实验,涵盖方法创新与实验设计,其评判标准泄漏率降至3.7%,而现有数据集的泄漏率为11.90%至34.10%。训练采用评分标准两次:首次作为OPSD自教师的特权上下文,随后作为GRPO的奖励。在Qwen3-1.7B/4B/8B模型上,该训练流程优于监督微调、任一单独阶段及反向顺序,使五项基准的平均得分分别提升5.6、5.0、4.8个百分点。采用相同方案时,在PaperGym-20k上训练的模型在三方对比中胜率达58.1%,而RubricHub Science仅为28.2%。训练后的Qwen3-8B在ResearchQA上得分为73.48,优于规模大得多的Kimi K2.6。本文发布了该流程、含20000个实例的语料库PaperGym-20k,以及基准PaperGym-Innov和PaperGym-Design。

英文摘要

Research planning is the decisive capability of AI scientists. Yet a research plan admits no verifiable answer, so reinforcement learning lacks the environment it requires: tasks paired with a critic. Rubrics extracted from scientific papers can supply the critic. Existing pipelines, however, draw the question and the criteria from the same content, so the reward can be earned by paraphrase. The rubric is further compressed into a single scalar per rollout. We introduce PaperGym, a unified framework that turns each research paper into a complete training environment. PaperGym exploits the structure of a paper: the question is synthesized from the research goal and background, while the criteria are derived from the method and experiments. The criteria span methodological innovation and experimental design, and criterion leakage falls to 3.7%, versus 11.90% to 34.10% in existing datasets. Training uses the rubric twice: first as privileged context for OPSD's self-teacher, then as the reward for GRPO. Across Qwen3-1.7B/4B/8B, this schedule outperforms supervised fine-tuning, either stage alone, and the reverse ordering, improving five-benchmark averages by +5.6, +5.0, and +4.8 points. With the recipe held fixed, models trained on PaperGym-20k win 58.1% of three-way comparisons, against 28.2% for RubricHub Science. The trained Qwen3-8B reaches 73.48 on ResearchQA, above the far larger Kimi K2.6. We release the pipeline, the 20,000-instance corpus PaperGym-20k, and the benchmarks PaperGym-Innov and PaperGym-Design.

Comments34 pages, 6 figures, 6 tables. Code: https://github.com/ZJU-REAL/PaperGym. Project page: https://zju-real.github.io/PaperGym. Dataset: https://huggingface.co/datasets/CabbageWyh/PaperGym-Data. Model: https://huggingface.co/CabbageWyh/PaperGym-Model

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑