SETA:终端智能体的扩展环境
SETA: Scaling Environments for Terminal Agents
浏览论文内容
中文总结 AI 辅助
研究针对终端智能体训练扩展难的问题,提出SETA框架,含SETA - Synth和SETA - Evol两个管道及统一验证机制,构建了SETA - Env数据集。实验显示该数据集能为终端智能体提供优质训练环境,推动相关研究发展。
中文摘要 AI 辅助
大语言模型正迅速向通过多种接口(包括网络和图形用户界面)解决任务的智能体转变。其中,终端命令行提供了基于文本的通用接口,涵盖从系统操作到数据科学和机器学习的任务。然而,扩展终端智能体训练仍具有挑战性,因为它需要多样且连贯的任务指令、可执行环境和可靠验证,同时缺乏自然基础的监督数据。在这项工作中,我们提出了SETA,这是一个用于为强化学习生成可验证终端环境的可扩展框架。该框架由两个共享统一验证机制的管道组成:SETA - Synth将各种来源转换为标准化的强化学习环境,SETA - Evol通过对难度和多样性的自适应控制从现有环境进一步扩展。我们共同构建并发布了SETA - Env,这是迄今为止最大的开源可验证终端强化学习数据集,包含超过4500个环境。我们通过在SETA - Env上使用GRPO训练Qwen3 - 8B来评估我们的数据集,在Terminal - Bench 2.0上实现了12%的通过率,这是8B规模的强化学习训练模型报告的最佳结果。我们还观察到在相同终端智能体框架下DeepSeek - V4 - Flash的收益,Terminal - Bench 2.0上的pass@1从40%提高到43%,pass@5从54%提高到58%。这些结果表明SETA - Env为终端智能体提供了高质量的训练环境,并作为推进基于终端的智能体学习研究的宝贵资源。
英文摘要
Large language models (LLMs) are rapidly shifting toward agents that solve tasks through diverse interfaces, including web and graphical user interfaces (GUIs). Among these, the terminal command line provides a text-based, general-purpose interface, covering tasks from system operations to data science and machine learning. However, scaling terminal-agent training remains challenging, as it requires diverse and coherent task instructions, executable environments, and reliable verification, while lacking naturally grounded supervision data. In this work, we propose SETA, a scalable framework for generating verifiable terminal environments for reinforcement learning (RL). The framework consists of two pipelines sharing a unified verification mechanism: SETA-Synth converts diverse sources into standardized RL environments, and SETA-Evol further expands from existing environments with adaptive control of difficulty and diversity. Together, we construct and release SETA-Env, the largest open-source verifiable terminal RL dataset to date, containing over 4,500 environments. We evaluate our dataset by training Qwen3-8B with GRPO on SETA-Env, achieving 12% pass rate on Terminal-Bench 2.0, the best reported result for an RL-trained model at the 8B scale. We further observe gains on DeepSeek-V4-Flash under the same terminal agent harness, with pass@1 on Terminal-Bench 2.0 improving from 40% to 43% and pass@5 improving from 54% to 58%. These results demonstrate that SETA- Env provides high-quality training environments for terminal agents and serves as a valuable resource for advancing research on terminal-based agent learning.
发表机构
- Imperial College London(帝国理工学院)
- University College London(伦敦大学学院)
- SambaNova(桑巴诺瓦公司)
- KAUST(阿卜杜拉国王科技大学)
- Stanford University(斯坦福大学)
- University of Oxford(牛津大学)
- University of Waterloo(滑铁卢大学)
- RadixArk(基数方舟公司)
机构由 AI 辅助整理,请以论文原文为准。