arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越人类监督扩展大型推理模型:通往超级智能的路径

Scaling Large Reasoning Models beyond Human Supervision: A Path toward Superintelligence

Zhiqin Yang, Jingwen Fu, Yuhan Liu, Hengyu Liu, Yonggang Zhang, Kainan Cao, Zizhuo Zhang, Chenxin Li, Ruibin Yuan, Jiahao Pan, Jiankai Sun, Zhenyuan Zhang, Yibo Li, Yunlong Lin, Jing Xiong, Sida Lin, Bo Han, Wei Xue, Yike Guo

arXiv 2608.31075首次发表:更新:

发表机构

The Hong Kong University of Science and Technology; Zhongguancun Academy; Xi’an Jiaotong University; The Chinese University of Hong Kong; The University of Hong Kong; Hong Kong Baptist University; Hunyuan Tencent; National University of Singapore; Xiamen University(香港科技大学; 中关村学院; 西安交通大学; 香港中文大学; 香港大学; 香港浸会大学; 腾讯混元; 新加坡国立大学; 厦门大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文研究人类监督逐步退出时大型推理模型(LRMs)的扩展路径,提出五级阶梯框架,分析自主化带来的风险,围绕策略能力等开展评估,为开发通往超级智能的自我维持学习系统提供结构化方案。

AI 中文摘要

近期大型推理模型(LRMs)的进展表明,带有可验证奖励的强化学习(RLVR)可显著提升数学与代码领域的推理能力,这些领域的结果可自动校验。将该进展扩展至开放式及智能体任务仍存在困难,因为可靠奖励更难获取,且直接人类监督无法跟上模型生成经验的规模与复杂度。本文研究当人类监督逐步退出学习循环时,LRMs如何继续提升性能。我们考察该问题的两个关联维度:奖励维度追踪从每实例人类判断到可复用验证器及奖励的发展,这类奖励甚至无需人类反馈即可运作;经验维度研究学习如何从人类策划的任务与环境,转向自主生成的课程、构建的环境及自主协同进化。我们通过从L0到L4的五级阶梯关联这两个维度,明确学习过程中哪些部分仍处于人类持续控制之下。我们的分析还强调了奖励与经验生成日益自主化带来的风险,包括奖励黑客、反馈漂移、课程崩溃及环境错误。因此,我们围绕三个互补对象提供评估:策略能力、反馈保真度及经验质量。该分析结构化阐述了超越人类监督扩展LRMs的现有方法,以及开发自我维持学习系统以通往超级智能过程中涉及的开放问题。此外,我们维护着一个持续更新的GitHub仓库,以跟踪最新进展。

英文摘要

Recent advances in large reasoning models (LRMs) have shown that reinforcement learning with verifiable rewards (RLVR) can substantially improve reasoning in mathematics and code, where outcomes can be checked automatically. Extending this progress to open-ended and agentic tasks remains difficult because reliable rewards are harder to obtain and direct human supervision cannot keep pace with the scale and complexity of model-generated experience. This paper studies how LRMs can continue to improve as human supervision gradually recedes from the learning loop. We examine two connected dimensions of this problem. The reward axis traces the development from per-instance human judgments to reusable verifiers and rewards that operate even without human feedback. The experience axis examines how learning can progress from human-curated tasks and environments toward self-generated curricula, constructed environments, and autonomous co-evolution. We connect these dimensions through a five-level ladder from L0 to L4 that identifies which parts of the learning process remain under continued human control. Our analysis further highlights the risks introduced by increasingly autonomous rewards and experience generation, including reward hacking, feedback drift, curriculum collapse, and environment errors. Consequently, we also provide the evaluation around three complementary objects: policy capability, feedback fidelity, and experience quality. This analysis provides a structured account of current approaches to scaling LRMs beyond human supervision and the open problems involved in developing self-sustaining learning systems toward superintelligence. Furthermore, we maintain a continuously updated \href{https://github.com/visitworld123/Awesome-Scaling-LRM-Beyond-Human-Supervision}{GitHub repository} to track the latest advances.

Comments72pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑