arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.05923cs.AI

VERA:面向智能体协同进化的可验证环境规模化

VERA: Scaling Verifiable Environments for Agentic co-Evolution

Junqi Liu, Yongyang Pan, Zhuosong Jiang, Dongbai Li, Bo Zhang, Xitong Ling, Sheng Wang, Hanrong Ye, Yufan He, Can Zhao, Pengfei Guo, Dong Yang, Andriy Myronenko… 展开作者

Junqi Liu, Yongyang Pan, Zhuosong Jiang, Dongbai Li, Bo Zhang, Xitong Ling, Sheng Wang, Hanrong Ye, Yufan He, Can Zhao, Pengfei Guo, Dong Yang, Andriy Myronenko, Yuyin Zhou, Tianyu Liu, Daguang Xu, Yucheng Tang

首次发表
浏览论文内容

中文总结 AI 辅助

提出VERA框架,通过可验证沙箱环境规模化构建与协同进化更新,使智能体在长时程任务中显著超越基线,并保持通用能力。

中文摘要 AI 辅助

胜任的智能体需要精确且可验证的环境,例如可在任何阶段恢复并能从可观察证据中演化的沙箱。然而,大多数长时程工作暴露了此类环境的稀缺性:例如,医学研究中的智能体必须在数十个依赖步骤中证实发现、对其进行分类并撰写报告,而近期环境仅对最终结果评分。为应对稳定训练中的挑战,我们提出VERA,它规模化构建此类环境并让智能体在其上演化。VERA从初始轨迹构建这些环境:智能体编写评分标准、可执行检查,一个评判者验证每个沙箱,只有通过的沙箱进入训练库。在这些环境上,VERA交替进行两种更新:用评分标准奖励训练模型,或编辑工具技能。我们还创建了一个验证器,使用明确的开发集验收标准来门控模型检查点和工具编辑。这种归因使VERA的协同进化区别于单轴基线:其更新不仅针对原因,还针对结果。凭借包含9000多个长时程可验证环境的开源语料库,一个9B模型与其协同进化的智能体配对,在两个领域中分别以10.3和13.0分的优势超越最强基线。在27B规模下,它在AutoCoWorkBench(71.6)和AutoMedBench(80.7)上超越基线,迁移到未见工作流,并保留通用能力。

英文摘要

Competent agents need precise and verifiable environments, such as sandboxes that are resumable at any stage and evolve from observable evidence. However, most long-horizon work exposes how rare these are: for example, an agent in medical research must ground a finding, classify it, and write a report over dozens of dependent steps, yet recent environments score only the outcome. To address the challenges in stable training, we present VERA, which builds such environments at scale and lets agents evolve on them. VERA builds these environments from initial trajectories: an agent writes rubrics, executable checks, a judge verifies each sandbox, and only those that pass enter the training bank. On these environments, VERA alternates between two updates: train the model with rubric rewards, or edit the harness skills. We also create a verifier which gates model checkpoints and harness edits using explicit development-set acceptance criteria. This attribution distinguishes VERA's co-evolution from single-axis baselines: its updates target not only the cause but the outcome. With an open-source corpus of 9,000+ long-horizon verifiable environments, a 9B model paired with its co-evolved agent beats the strongest baseline by 10.3 and 13.0 points in the two domains. At 27B, it surpasses the baseline on AutoCoWorkBench (71.6) and AutoMedBench (80.7), transfers to unseen workflows, and retains general capabilities.

发表机构

  • Nvidia(英伟达)
  • University of California, Santa Cruz(加州大学圣克鲁兹分校)
  • University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
  • National University of Singapore(新加坡国立大学)
  • Tsinghua University(清华大学)
  • Xian Jiaotong University(西安交通大学)
  • University of Pennsylvania(宾夕法尼亚大学)

机构由 AI 辅助整理,请以论文原文为准。

↑