arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

NeoHorse-1:通过带路由框架的智能体后训练实现递归自我改进

NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness

NeoHorse Team, Guoliang Cao, Guohao Dai, Tianyu Guo, Kai Han, Hailin Hu, Zihan Jiang, Xiang Kuang, Boxun Li, Yulong Li, Zehua Pei, Yuchuan Tian, Jiamin Wang, Yu Wang, Yunhe Wang, Yihong Wu, Haiyang Xu, Shuo Zhang, Hang Zhou, Siyang Cheng, Jiayu Fan, Wei He, Qingrui Jiao, Hongguang Li, Zhiyuan Li, Runke Liu, Xi Liu, Xinchen Liu, Sinno Jialin Pan, Yi Ren, Liuyang Song, Chenyu Wang, Bei Yu, Quanlu Zhang, Xiangyu Zhang, Mengyu Zheng, Yingjie Zong

arXiv 2609.08183首次发表:更新:

AI 中文总结

NeoHorse-1通过智能体后训练与路由框架,记录交互并转化为训练数据,实现递归自我改进,在11个基准上显著提升模型性能。

AI 中文摘要

递归自我改进(RSI)需要一个具体机制,使人工智能系统能够观察自身能力并将该证据转化为下一轮学习。我们提出了NeoHorse-1,一个通过智能体后训练探索这一路径的智能体原生模型系列。我们的系统将异构模型池与智能路由相结合,记录每个用户回合的预测能力需求、所选服务层级及后续交互。这些记录被转化为保留交错推理、工具调用和框架上下文的训练示例,并通过结构验证、六维语义评估和子场景级标注进行筛选。路由信号将监督微调组织为三阶段课程,并扩展到路由引导的在线策略蒸馏,其中教师在同一进度下监督学生生成的响应。能力引导分配随后将评估反馈转化为下一训练混合,形成一个评估-选择-更新循环,系统学习做什么决定了它下一步从什么中学习。在涵盖基于框架的智能体、工具使用、编码和指令遵循的十一个基准上,后训练将宏观平均值从4B规模的58.94提升至64.87,从9B规模的65.60提升至69.04,大幅缩小了后训练4B模型与9B基础模型之间的总体差距。NeoHorse-1提供了这一反馈驱动过程的初始原型,以及跨连续迭代实现框架介导的RSI的路径。

英文摘要

Recursive self-improvement (RSI) requires a concrete mechanism through which an AI system observes its capabilities and converts that evidence into the next round of learning. We present NeoHorse-1, a family of agent-native models developed to explore this path through agentic post-training. Our system combines a heterogeneous model pool with intelligent routing, recording the predicted capability demand, selected service tier, and subsequent interaction for each user turn. These records are converted into training examples that preserve interleaved reasoning, tool calls, and harness context, and are admitted through structural validation, six-dimensional semantic evaluation, and subscene-level labeling. Routing signals organize supervised fine-tuning into a three-stage curriculum and extend to routing-guided on-policy distillation, where a teacher supervises student-generated responses under the same progression. Capability-guided allocation then converts evaluation feedback into the next training mixture, closing an evaluation-selection-update loop in which what the system learns to do shapes what it learns from next. Across eleven benchmarks covering harness-based agents, tool use, coding, and instruction following, post-training raises the macro-average from 58.94 to 64.87 at 4B and from 65.60 to 69.04 at 9B, substantially narrowing the aggregate gap between the post-trained 4B model and the 9B base model. NeoHorse-1 provides an initial prototype of this feedback-driven process and a path toward harness-mediated RSI across successive iterations.

CommentsHuggingface: https://hf.co/collections/TokenRhythm/neohorse-1; Github: https://github.com/TokenRhythm/NeoHorse

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑