ToolRACER:用于智能体训练与评估的稳健代理式对话模拟资源
ToolRACER: A Robust Agentic Conversation Emulation Resource for Agent Training and Evaluation
浏览论文内容
中文总结 AI 辅助
ToolRACER提出合成数据生成流水线,构建多轮对话基准ToolRACERBench,注入对抗行为,训练模型提升智能体在函数调用基准上的端到端准确性与稳健性。
中文摘要 AI 辅助
面向任务的对话智能体在真实世界对话场景中仍然表现脆弱,因为它们很少遵循可预测的脚本,尤其是当用户表现出不合作行为时。现有的函数调用基准通常强调成功的、合作性的交互,而低估了对抗性对话轨迹,从而限制了可用于开发稳健智能体的训练资源。我们提出了ToolRACER,一个合成数据生成流水线,它协调用户、助手和工具模拟模型来生成并验证用户与智能体之间的多轮交互。使用ToolRACER,我们构建了ToolRACERBench,一个涵盖六个领域、超过55种不同人物角色的稳健多轮对话基准,生成了包含5.6K条对话轨迹的验证语料库,其中约66%的对话包含易失败的对话场景。我们注入对抗性行为,产生验证过的对话交互轨迹,以捕捉真实、稳健的场景。我们评估了在ToolRACERBench上训练的模型,与内部基准以及函数调用基准(如τ²-bench、BFCLv3和ACEBench)进行比较,以评估智能体的准确性和稳健性。在ToolRACERBench上训练的模型在τ²-bench和ACEBench上提高了端到端的智能体准确性,表明在与小语言模型中的域内数据集混合时,对智能体能力任务有显著提升。
英文摘要
Task-oriented conversational agents remain fragile under real world conversation scenarios as they rarely follow a predictable script, especially when users exhibit non-cooperative behavior. Existing function-calling benchmarks often emphasize successful, cooperative interactions and underrepresent adversarial conversation trajectories, thereby limiting the training resources available for developing robust agents. We present ToolRACER, a synthetic data generation pipeline that coordinates user, assistant and tool emulation models to generate and validated multi-turn interactions between a user and an agent. Using \sysn, we construct ToolRACERBench a robust multi-turn conversation benchmark spanning six domains, ranging over 55 varied personas, generating a validated corpus of 5.6K conversation trajectories, with approximately 66\% of conversations containing failure-prone conversation scenarios. We inject adversarial behaviors, producing validated conversational interaction trajectories that capture realistic, robust scenarios. We evaluate models trained on ToolRACERBench against internal benchmarks, as well as on function calling benchmarks such as $τ^2$-bench, BFCLv3 and ACEBench to evaluate agentic accuracy and robustness. Models trained on ToolRACERBench improve end to end agentic accuracy across $τ^2$-bench and ACEBench, demonstrating significant gains when mixed with in-domain dataset in small language models for agent capability tasks.
发表机构
- Uniphore
- University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
机构由 AI 辅助整理,请以论文原文为准。