Terminal-Universe:将智能体轨迹转化为可扩展的终端环境
Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments
- Qwen Team, Alibaba Group(通义千问团队,阿里巴巴集团)
- Tsinghua University(清华大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
Terminal-Universe 是将智能体轨迹转化为可重用终端环境的框架,可扩展任务的广度与深度,经其生成的语料库微调 Qwen3.5-27B,显著提升了代码智能体的单轮与多轮任务性能。
AI中文摘要:
随着基于终端的代码智能体变得普及,智能体轨迹已大规模积累,但逼真的可执行环境仍然稀缺。然而,环境是智能体后训练实际所需的:每个环境可被重查询为许多可验证任务并提供执行反馈,而轨迹是单个冻结的演示。我们观察到,现有轨迹中的工具执行历史暴露了其运行环境的结构和内容,无需从头生成环境,即可从轨迹本身重建这些环境。因此,我们提出Terminal-Universe,这一框架将每个轨迹转化为可重用环境,并探索其以合成新任务和持续交互。具体而言,Terminal-Universe 重放轨迹中记录的文件操作,以恢复智能体修改前的每个文件,生成部分工作空间;随后一个补全智能体提供缺失的文件和依赖项。在这个恢复的工作空间上,我们既重建原始意图任务,也合成全新任务。此外,我们还沿两个互补轴扩展任务:广度和深度。广度方面,我们挖掘相关环境间的定向依赖关系,合成跨多个代码库的跨工作空间查询,如同开发者在实际开发中常做的那样。深度方面,我们将初始单轮查询扩展为多轮会话,通过用户智能体捕捉迭代的用户反馈和需求细化。应用于公开终端智能体轨迹,Terminal-Universe 生成了3.73万个任务充足的环境。在该语料库上对 Qwen3.5-27B 进行监督微调,使 Terminal-Bench 2.1 上的单轮性能提升11.9个百分点,EvoCode-Bench v2 MT@4 上的多轮性能提升13.8个百分点。
英文摘要:
As terminal-based code agents become prevalent, agent trajectories have accumulated at scale, while realistic, executable environments remain scarce. However, environments are what agent post-training actually requires: each can be re-queried into many verifiable tasks and provides execution feedback, whereas a trajectory is a single frozen demonstration. Rather than generating environments from scratch, we observe that the tool-execution history in existing trajectories exposes the structure and contents of the environments in which they ran, making it possible to reconstruct those environments from the trajectories themselves. Thus, we introduce Terminal-Universe, a framework which turns each trajectory into a reusable environment and explores it for synthesizing new tasks and continued interactions. Specifically, Terminal-Universe replays the file operations recorded in a trajectory to restore each file before the agent modified it, yielding a partial workspace; a completion agent then supplies the missing files and dependencies. On this recovered workspace, we both reconstruct the original intent task and synthesize entirely new ones. Besides, we also scale the tasks along two complementary axes: breadth and depth. For breadth, we mine directional dependency relations between related environments and synthesize cross-workspace queries spanning multiple codebases, as developers routinely do in real-world development. For depth, we extend the initial single-turn query into a multi-round session that captures iterative user feedback and requirement refinement via a user agent. Applied to public terminal agent trajectories, Terminal-Universe produces 37.3k task-sufficient environments. Supervised fine-tuning of Qwen3.5-27B on this corpus improves single-round performance on Terminal-Bench 2.1 by 11.9 points and multi-round performance on EvoCode-Bench v2 MT@4 by 13.8 points.