arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RecreationWorld:面向混合型计算机使用智能体的可扩展且可验证环境

RecreationWorld: Scalable and Verifiable Environments for Hybrid Computer-Use Agents

Shuai Bai, Jiayong Deng, Sicheng Fan, Yikun Fu, Chang Gao, Xuhao Hu, Mianqiu Huang, Yizhen Jiang, Yuheng Jing, Dehui Kong, Keliang Li, Ning Li, Wanli Li, Dayiheng Liu, Dunjie Lu, Changwei Luo, Que Shen, Zheyuan Wang, Zijian Wang, Jie Wu, Gao Wu, Zhihui Xie, Rui Xie, Haiyang Xu, An Yang, Jiakang Yuan, Yanming Zhang, Jiajun Zhang, Xi Zhang, Zhenru Zhang, Zhuo Zhen, Mingkang Zhu, Bowen Zhou

arXiv 2609.22000首次发表:更新:

发表机构

Alibaba Token Hub; Alibaba Group(阿里巴巴通义实验室; 阿里巴巴集团)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出RecreationWorld,一个五平台可扩展环境,用于训练和评估混合计算机使用智能体,通过复现参考应用行为并提供可验证测试,显著提升跨域迁移能力。

AI 中文摘要

计算机使用智能体(CUA)沿着两条独立的路线发展:图形界面交互,以及通过代码和命令行进行软件开发。真实的数字工作两者都需要,且需要交错进行,而非首尾简单叠加。我们研究混合型CUA,它们自主决定何时探索界面、实现软件,以及运行并视觉验证其产物。我们引入RecreationWorld,一个围绕“复现”构建的五平台框架:给定一个运行中的参考实现,智能体必须发现其行为,并在没有规定工作流程的情况下构建忠实的实现。RecreationWorld在Ubuntu、macOS、Windows、Android和Web上提供可复现的环境,并配备统一测试平台,支持原生GUI控制和编码工具。运行中的参考实现充当隐藏行为测试的预言机,提供基于执行结果的奖励。我们利用高质量的开源应用来扩展轨迹生成。在这些轨迹上训练的模型在五个分布外编码和混合计算机使用基准上取得改进,并更频繁地验证其渲染输出,为超越复现的迁移提供了证据。对于留出评估,我们引入RecreationBench,包含跨领域和平台的250个多样化任务。基于参考的程序化和视觉断言覆盖了多个交互深度下的动作条件结果;每个断言都在参考实现上验证,并由人工审查,之后冻结套件用于自动评分。GPT-6 Astra以58.1%的整体通过率领先,但仅在2.8%的任务上通过所有程序化测试。智能体在复现静态界面结构方面比交互和计算输出更可靠,而生成的应用程序仍比参考实现更小、更单体化。我们发布基准、环境和测试套件。

英文摘要

Computer-use agents (CUAs) have advanced along two separate lines: graphical interaction and software development through code and the command line. Real digital work requires both, interleaved rather than stacked end to end. We study hybrid CUAs that autonomously decide when to explore an interface, implement software, and run and visually verify their artifacts. We introduce RecreationWorld, a five-platform framework built around recreation: given a running reference, an agent must discover its behavior and build a faithful implementation with no prescribed workflow. RecreationWorld provides reproducible environments on Ubuntu, macOS, Windows, Android, and Web, plus a unified harness with native GUI control and coding tools. The running reference serves as an oracle for hidden behavioral tests, providing execution-grounded rewards. We scale trajectory generation with high-quality open-source applications. Models trained on these trajectories improve across five out-of-distribution coding and hybrid computer-use benchmarks and more frequently verify their rendered outputs, providing evidence of transfer beyond recreation. For held-out evaluation, we introduce RecreationBench, comprising 250 diverse tasks across domains and platforms. Reference-grounded programmatic and visual assertions cover action-conditioned outcomes at multiple interaction depths; each is validated on the reference and by human reviewers before the suite is frozen for automatic scoring. GPT-6 Astra leads at 58.1% overall, but passes all programmatic tests on just 2.8% of tasks. Agents reproduce static interface structure more reliably than interactions and computed outputs, while generated applications remain smaller and more monolithic than their references. We release the benchmark, environments, and test suites.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑