arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Echoverse:用于大规模训练计算机使用智能体的深度演化环境

Echoverse: Deep, Evolving Environments for Training Computer-Use Agents at Scale

Yash Pandya, Sahil Gupta, Sarthak Harne, Archana Yadav, Kavyansh Chourasia, Hussein Mozannar, Vibhav Vineet, Sara Abdali, Corby Rosset, Yash Lara, Ahmed Awadallah, Ece Kamar, Akshay Nambi

arXiv 2607.28074首次发表:更新:

发表机构

Microsoft Research(微软研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Echoverse是用于大规模训练计算机使用智能体的深度演化环境,其协同演化循环可提升智能体性能,经12个环境训练后9B模型准确率大幅提升,还发布了基准环境与代码

AI 中文摘要

计算机使用智能体从自身动作带来的变化中学习,因此训练这类智能体需要可供其执行、破坏和重置的应用程序。最重要的应用程序受登录限制且具有状态,因此合成环境可作为它们的替代。近期的流程能够批量生成这类环境,这使得瓶颈从环境数量转向每个环境的内部质量。我们发现,收益来自三个特性:环境具备的行为深度、是否针对智能体实际会失败的交互、是否能随模型一同改进。我们提出Echoverse,它能将规范编译为有状态应用,其任务依据应用自身的数据库进行分级;还设计了协同演化循环,会对每个分级的 rollout 进行两次处理:一次作为环境、任务和验证器的修复,另一次作为模型的训练信号。在12个此类环境上训练后,一个9B模型在14个评估拆分中的准确率从36.5%提升至67.1%,与指导它的更大前沿模型仅相差14个百分点。我们逐一检验每个特性:在相同领域,浅环境会将现场准确率降至基础模型以下(80.0%→75.0%),而深环境会提升准确率(80.0%→85.0%、48.0%→65.0%);对多个渲染版本的同一界面控制进行训练,可泛化到未见过的控件家族和开放网页;修复单个环境可使基于其训练的模型从16.2%提升至38.5%。这些相同的环境还可作为强化学习环境,结合基于接地验证器与逐步密集评判器的奖励,将未见过的得分从58.8%提升至68.0%。我们发布了4个环境作为基准,附带其应用程序、种子数据和接地评判器。代码:this https URL

英文摘要

Computer-use agents learn from what their actions change, so training one needs applications it can act on, break and reset. The applications that matter most are login-gated and stateful, so synthetic environments stand in for them. Recent pipelines generate such environments in bulk, which moves the bottleneck from how many exist to what is inside each one. The returns, we find, come from three properties: how much behavioural depth an environment carries, whether it targets the interaction an agent actually fails, and whether it improves alongside the model. We present Echoverse, which compiles specifications into stateful applications whose tasks are graded against the application's own database, and a co-evolution loop that reads every graded rollout twice: as repairs to the environment, its tasks and its verifier, and as training signal for the model. Trained on twelve such environments, a 9B model improves from $36.5\%$ to $67.1\%$ across fourteen evaluation splits, within fourteen points of the much larger frontier model that taught it. We examine each property in turn. On the same domains, shallow environments push live-site accuracy below the base model ($80.0 \to 75.0$) while deep ones raise it ($80.0 \to 85.0$ and $48.0 \to 65.0$); drilling one interface control across many renderings transfers to held-out widget families and to the open web; and repairing a single environment lifts the model trained on it from $16.2\%$ to $38.5\%$. The same worlds serve as reinforcement-learning environments, where a reward combining the grounded verifier with a dense per-step judge raises held-out score from $58.8\%$ to $68.0\%$. We release four environments as a benchmark, with their applications, seed data and grounded graders. Code: https://aka.ms/echoverse

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑