WorkGenesis:构建教会智能体工作的世界
WorkGenesis: Building the Worlds That Teach Agents to Work
- Shanghai Jiao Tong University(上海交通大学)
- Endless Frontier
- Beihang University(北京航空航天大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
WorkGenesis通过基于证据的工作构建和执行引导的一致性验证,从现实工件合成可执行职业工作,为训练工作智能体提供可扩展数据,其训练的Fx-Work-35B在多个基准上超越更大规模模型。
AI中文摘要:
大型语言模型(LLM)智能体完成日常工作和专业工作的能力正受到越来越多的关注。训练此类智能体需要逼真的工作场景。专家撰写的职业工作数据成本高昂且制作缓慢,而无约束的合成往往会产生事实依据薄弱或内部要求不一致的任务。为弥合这一差距,我们提出了WorkGenesis,一个通过两项核心技术创新从现实世界工件构建可执行职业工作的框架:(1)基于证据的工作构建,该方法通过O*NET职业知识检索公开文件,并围绕这些文件合成上下文、配套材料、工作请求和逐项评分标准,从而将每个工作单元锚定在现实世界的证据上;(2)执行引导的一致性验证,该方法在构建的工作内部渲染参考交付物,将每个未满足的评分标准项归因于智能体、任务或评分标准,并利用任务和评分标准的缺陷作为反馈,迭代修复工作直至通过审计。实验结果表明,仅使用WorkGenesis合成的20K个工作单元进行简单监督微调(SFT)训练的Fx-Work-35B,在GDPvalAA-v2、APEX-Agents-AA和JobBench上报告的五项指标中,在所有可比规模的基线中取得了最高分数(平均得分31.00对比24.79),甚至超越了诸如1.6T DeepSeek-V4-Pro-Preview等前沿模型。这些结果表明,WorkGenesis为工作智能体提供了可扩展的训练数据。
英文摘要:
The ability of Large Language Model (LLM) agents to complete daily and professional work is receiving increasing attention. Training such agents requires realistic work scenarios. Expert-authored occupational work is costly and slow to produce, while unconstrained synthesis often yields tasks with weak factual grounding or internally inconsistent requirements. To bridge this gap, we introduce WorkGenesis, a framework that constructs executable occupational work from real-world artifacts through two core technical innovations: (1) Evidence-Based Work Construction, which grounds each unit of work in real-world evidence by retrieving public files guided by O*NET occupational knowledge and synthesizing the surrounding context, companion materials, work request, and itemwise rubric around them; and (2) Execution-Guided Consistency Verification, which renders a reference deliverable inside the constructed work, attributes every unsatisfied rubric item to the agent, the task, or the rubric, and uses task and rubric defects as feedback to iteratively repair the work until it passes the audit. Experimental results demonstrate that Fx-Work-35B, trained with simple supervised fine-tuning (SFT) on only 20K units of work synthesized by WorkGenesis, achieves the highest scores among all comparable-scale baselines on the five reported metrics across GDPvalAA-v2, APEX-Agents-AA, and JobBench (31.00 versus 24.79 average score), and even surpasses frontier models such as the 1.6T DeepSeek-V4-Pro-Preview. These results show that WorkGenesis provides scalable training data for working agents.