arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38923cs.CL

GraphForge:基于图锚定工作区合成训练工作型智能体

GraphForge: Training Working Agents with Graph-Anchored Workspace Synthesis

Qisheng Su, Hanchen Wang, Guanru Zhu, Huicheng Jiang, Qiuyinzhe Zhang, Kou Shi, Zhen Fang, Ziao Zhang, Qingnan Ren, Honglin Guo, Zehui Chen, Tao Gui, Feng Zhao

首次发表
浏览论文内容

中文总结 AI 辅助

GraphForge通过证据图将任务与验证锚定在真实文件上,合成工作型智能体训练数据,微调Qwen3.6-27B显著提升多个基准性能。

中文摘要 AI 辅助

工作型智能体需要读取多种文件、协调工具并生成交付物。训练此类智能体需要基于大量真实文件且结果可验证的任务,但现有流程很少能合成这类数据。现有流程要么用模型生成文件,缺乏真实性和多样性;要么在真实文件上构建任务但缺乏针对任务的验证器,导致结果质量无法检查。我们提出GraphForge,一个基于证据图的框架,将任务及其验证都锚定在真实文件中。从基于职业的种子出发以实现受控多样性,GraphForge为每个种子组装一个由真实文件组成的工作区,并在其关系上构建证据图。由于任务陈述和评分标准均源自该图,任务要求由工作区文件支撑,每条标准都锚定在验证所需文件上。初始试运行进一步测试可执行性,修订智能体在收集轨迹前根据原始文件修复任务和评分标准。在2,169条GraphForge轨迹上微调Qwen3.6-27B,在OpenHands下将GDPVal提升至1445.7(+65.7),在Claude Code下将Workspace-Bench-Lite和SpreadsheetBench II分别提升至63.7(+7.7)和24.0(+13.7)。对SFT模型自身试运行进行拒绝微调,候选由证据锚定评分标准选择,在三个基准上均取得进一步改进,表明评分标准提供了有用的选择信号。数据和模型已公开。

英文摘要

Working agents need to read diverse files, coordinate tools, and produce deliverables. Training such agents requires tasks built on many real files with verifiable results, but few pipelines exist to synthesize this kind of data. Existing pipelines either generate files with models, which lack realism and diversity, or build tasks on real files without task-specific verifiers, leaving result quality unchecked. We introduce GraphForge, an evidence-graph based framework that grounds both the task and its verification in real files. Starting from occupation-grounded seeds for controlled diversity, GraphForge assembles a workspace of real files for each seed and builds an evidence graph over their relations. Since the task statement and rubrics are both derived from this graph, task requirements are backed by the workspace files and each criterion is anchored to the files needed to verify it. An initial rollout further tests executability, and a revision agent repairs the task and rubrics against the original files before trajectories are collected. Fine-tuning Qwen3.6-27B on 2,169 GraphForge trajectories brings GDPVal to 1445.7 (+65.7) under OpenHands, and Workspace-Bench-Lite and SpreadsheetBench II to 63.7 (+7.7) and 24.0 (+13.7) under Claude Code. Rejection fine-tuning on the SFT model's own rollouts, with candidates selected by the evidence-anchored rubrics, yields further improvements on all three benchmarks, suggesting that the rubrics provide a useful selection signal.

发表机构

  • University of Science and Technology of China(中国科学技术大学)
  • Shanghai Innovation Institute(上海创新研究院)
  • Fudan University(复旦大学)
  • Shanghai AI Laboratory(上海人工智能实验室)

机构由 AI 辅助整理,请以论文原文为准。

↑