发表机构
SAP(SAP)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出EdgeGen框架,从智能体规范提取合规规则生成边界任务,结合微调与工具优化,实现无人工标注的闭环改进,在tau2bench上平均进度提升2%-42%。
AI 中文摘要
工具调用的大语言模型(LLM)智能体正越来越多地部署在企业应用中。然而,有效的评估和优化需要高质量、多样化的任务数据集,这些数据集往往因隐私和其他限制而难以获取。现有的合成任务生成方法通常产生通用任务,这些任务忽略了智能体的底层状态或数据库,未能反映真实世界的使用多样性。我们提出了EdgeGen,一个合成任务生成框架,它从智能体的规范中提取合规规则,并利用这些规则生成基于数据库的边界情况任务,这些任务旨在违反这些规则。当与现有的合成数据生成技术结合时,EdgeGen通过微调和工具(harness)优化实现智能体改进。由此产生的流程形成了一个完全自动化的闭环系统,无需人工标注。在tau2bench航空领域上,使用EdgeGen生成的数据进行微调,平均进度提升了一致性的2%至42%,而其他基线方法在某些模型上表现出性能下降。另一方面,对于工具优化,我们的方法在Gemma-4-e4b模型上,相较于人工策划的工具和基础工具,分别显示了10%和30%的平均进度提升。
英文摘要
Tool-calling LLM agents are increasingly deployed in enterprise applications. However, effective evaluation and optimization require high-quality, diverse task datasets that are often difficult to obtain due to privacy and other constraints. Existing synthetic task generation methods often produce generic tasks that ignore an agent's underlying state or database and fail to reflect real-world usage diversity. We propose EdgeGen, a synthetic task generation framework that extracts compliance rules from an agent's specification and uses them to generate database-grounded edge-case tasks designed to violate these rules. When combined with existing synthetic data generation techniques, EdgeGen enables agent improvement through finetuning and harness optimization. The resulting pipeline forms a fully automated closed-loop system that requires no human annotation. Finetuning on data generated by EdgeGen yields a consistent mean progress improvement of 2 percent to 42 percent on tau2bench airline domain, while other baseline methods show degradation for some models. On the other hand, for harness optimization, our method shows a mean progress improvement of 10 percent and 30 percent over the human-curated and base harnesses, respectively, for the Gemma-4-e4b model.
CommentsNA