用于 API 调用智能体的无环境合成数据生成
Simulate to Generalize: Scaling Stateful Supervision for API-calling Agents using LLM World Models
浏览论文内容
中文总结 AI 辅助
研究针对训练 API 调用 LLM 智能体数据收集瓶颈问题,提出无环境合成数据生成方法,利用 LLMs 生成任务、响应等,经评判器过滤轨迹,在相关基准上评估,结果表明此方法可有效训练智能体,是实用可扩展方案。
中文摘要 AI 辅助
训练 API 调用大型语言模型(LLM)智能体需要大量高质量轨迹。但大规模收集此类数据通常需要具有可执行 API 和预填充后端数据库的完整环境,这成为扩展性的主要瓶颈。为克服此问题,我们提出一种无环境合成数据生成方法,利用 LLMs 作为即时数字世界模型。仅给定 API 规范,该方法生成模仿智能体与有状态环境交互的轨迹。具体而言,LLM 先生成可用所提供 API 解决的各种任务,教师智能体迭代解决每个任务,同时 LLM 模拟器根据任务上下文和模拟历史生成连贯的合成 API 响应,最后 LLM 评判器过滤轨迹以确保数据集质量。我们在具有挑战性的 AppWorld 和 OfficeBench 基准上评估该方法,在合成数据上微调模型带来显著性能提升,表明无需任何可执行环境就能为 API 调用智能体生成有效监督。我们的结果确立了基于 LLM 的 API 模拟作为跨不同 API 生态系统训练智能体的实用、可扩展解决方案。
英文摘要
Training agents that generalize to unseen, stateful environments requires a massive dataset of state-changing trajectories covering a vast and diverse set of APIs. However, scaling this broad supervision is severely bottlenecked by the immense effort required to implement and populate fully-executable environments across a broad spectrum of domains. To bypass this barrier, we introduce a data generation pipeline that decouples data synthesis from environment construction by leveraging LLMs as digital world models. Starting from only a list of broad domain names, our automated pipeline synthesizes diverse APIs and tasks. To produce trajectories, a teacher agent iteratively solves these tasks while an LLM simulator dynamically tracks state and provides coherent API responses on-the-fly. Finally, an automated judge filters the trajectories for quality. Fine-tuning on our broad synthetic dataset yields significant performance gains on AppWorld and OfficeBench, two challenging stateful benchmarks featuring environments completely unseen during training. These results establish our LLM world model-based synthesis approach as a highly scalable path for training generalizable, stateful API-calling agents.