发表机构
Tencent; Fudan University(腾讯; 复旦大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出WorkForge框架,从真实工作流构建可验证的工作智能体环境,通过事实锚点生成任务与验证器,在40个领域构建16.7K环境,显著提升模型性能并展现扩展规律。
AI 中文摘要
工作智能体在数字工件上操作以执行专业的知识密集型工作,这要求训练环境支持长时程交互和可信的验证。然而,手工构建的环境会带来高昂的工程开销,阻碍了环境的规模化扩展,而合成方法则牺牲了工作空间的复杂性、真实性或基于事实的可验证性。为弥合这一差距,我们提出了WorkForge,一个可扩展的合成框架,用于从真实世界资源构建可验证的工作智能体环境。从专家工作流出发,WorkForge首先识别每个工作流所需的资源、决策和交付物。然后,它检索相关的真实世界文件并将其组织到工作空间中。WorkForge检查工作空间以提取关于其内容的具体的、可核查的事实。这些事实锚点确定了工作空间能够支持的任务类型以及如何验证其结果。因此,WorkForge直接从这些事实锚点推导出每个任务的指令、解决方案计划以及互补的程序化和语义验证器,使验证可追溯到可观察的工作空间证据。此外,我们构建了跨越40个专业领域的16.7K个可验证环境,工作空间共涵盖60种文件类型。后训练的Qwen3.5-35B-A3B-Base将GDPVal从45.5提升至73.6,APEX分数从5.0提升至21.3,同时使Qwen3.5-27B达到极具竞争力的性能并超越强竞争对手。我们的分析证实了所提方法的有效性,并揭示了在数据量和交互时程上的一致扩展行为。
英文摘要
Work agents operate over digital artifacts to execute professional knowledge-intensive work, requiring training environments that support long-horizon interaction and trustworthy verification. However, hand-crafted environments incur prohibitive engineering overhead that prevents environment scaling, whereas synthesis methods sacrifice workspace complexity, realism, or grounded verifiability. To bridge this gap, we introduce WorkForge, a scalable synthesis framework for constructing verifiable work-agent environments from real-world resources. Starting from expert workflows, WorkForge first identifies the resources, decisions, and deliverables required by each workflow. It then retrieves relevant real-world files and organizes them into a workspace. WorkForge inspects the workspace to extract concrete, checkable facts about its content. These factual anchors fix which task types the workspace can support and how their outcomes can be verified. Therefore, WorkForge derives each task's instructions, solution plan, and complementary programmatic and semantic verifiers directly from these factual anchors, keeping verification traceable to observable workspace evidence. Furthermore, we construct 16.7K verifiable environments across 40 professional domains, with workspaces collectively covering 60 file types. Post-training Qwen3.5-35B-A3B-Base improves GDPVal from 45.5 to 73.6 and APEX Score from 5.0 to 21.3, while enabling Qwen3.5-27B to achieve highly competitive performance and outperform strong competitors. Our analyses confirm the efficacy of the proposed method and reveal consistent scaling behaviors across both data volume and interaction horizons.
CommentsThis version was submitted before all co-authors had completed their review and approved the manuscript and author list. We are withdrawing it while these issues are resolved