arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.29175cs.SE

面向LLM智能体的优先执行式合成工具使用轨迹生成

Execution-First Synthetic Tool-Use Trace Generation for LLM Agents

Hafsa Ouajdi, Francesco Giannuzzo, Alaa Boukhary, Paolo Papotti, Gerard Conangla, Adam Elwood

首次发表
浏览论文内容

中文总结 AI 辅助

针对传统数据合成的局限,提出优先执行式框架SyntheticAgentTraceQA生成工具增强型智能体的监督数据,经四工具生态系统及Qwen模型验证,该框架可提升工具执行等性能,且揭示了掩码与全监督的权衡。

中文摘要 AI 辅助

智能体软件工程及工业系统日益通过可执行工作流而非仅代码生成运行:它们搜索制品、调用工具、检查结构化观测结果并查询数据库。训练这类智能体需要包含有效工具交互和可执行工作流的监督数据,但传统的优先查询式数据合成可能失效,因为看似合理的用户请求可能不对应有效工具序列、兼容参数或可用数据。为解决此局限,我们提出SyntheticAgentTraceQA,这是一种为工具增强型智能体生成可扩展监督数据的优先执行式框架。该框架首先构建高层工作流结构,通过感知依赖的分配将其映射到可用工具,在受控环境中执行并验证生成的轨迹,之后才合成自然语言用户任务、教师生成的推理注释及参考答案。我们在四个工具生态系统中评估该框架,并使用生成的数据微调和评估Qwen模型变体。结果表明,基于执行的监督可提升工具执行行为、参考轨迹一致性及评估任务的答案生成性能。进一步分析揭示了一种监督权衡:掩码监督(从训练目标中排除推理注释)可提升最终答案指标,而全监督(计算包含推理标记的完整助手输出的损失)在答案质量上表现不佳,且无法始终提升参考轨迹一致性,尤其在9B规模模型上。这些发现凸显了根据工具增强型智能体的期望能力设计合成监督的重要性。

英文摘要

Agentic software-engineering and industrial systems increasingly operate through executable workflows rather than code genera- tion alone: they search artifacts, invoke tools, inspect structured observations, and query databases. Training these agents requires supervision data that captures valid tool interactions and executable workflows. However, traditional query-first data synthesis can fail because plausible user requests may not correspond to valid tool sequences, compatible parameters, or available data. To address this limitation, we propose SyntheticAgentTraceQA, an execution- first framework for generating scalable supervision data for tool- augmented agents. Our framework first constructs high-level work- flow structures, maps them to available tools through dependency- aware assignment, executes and validates the resulting traces in con- trolled environments, and only then synthesizes natural-language user tasks, teacher-generated reasoning annotations, and reference answers. We evaluate the framework across four tool ecosystems and use the resulting data to fine-tune and evaluate Qwen model variants. The results show that execution-grounded supervision improves tool execution behavior, reference-trace agreement, and answer-generation performance on the evaluated tasks. Further analysis reveals a supervision trade-off: masked supervision, which excludes reasoning annotations from the training objective, im- proves final-answer metrics, whereas full supervision, computing loss over the complete assistant output including reasoning tokens, underperforms on answer quality and does not consistently im- prove reference-trace agreement, particularly at the 9B scale. These findings highlight the importance of designing synthetic supervi- sion according to the desired capabilities of tool-augmented agents.

发表机构

  • EURECOM
  • Aily Labs

机构由 AI 辅助整理,请以论文原文为准。

↑