arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.26546cs.AIcs.CL

DuMateBench:评估自主智能体在复杂真实工作流程中的表现

DuMateBench: Evaluating Autonomous Agents in Complex Real-World Workflows

Zechun Niu, Yukun Zhao, Jiaxin Zhang, Xu Shen, Jinhua Si, Han Tian, Can Xu, Yunfan Song, Jiaxin Mao, Yansong Gao, Yuchen Li, Jianmin Wu, Lingyong Yan, Shuaiqiang Wang, Dawei Yin

首次发表
浏览论文内容

中文总结 AI 辅助

DuMateBench是基于真实用户会话构建的自主智能体评估基准,含200个跨8场景的任务,注入三类环境复杂性,搭配五款智能体框架与四款LLM实验,发现严格任务完成存在显著差距。

中文摘要 AI 辅助

自主智能体正越来越多地被用于在真实场景中完成复杂的多工具工作流程。然而,现有的基准通常按应用或能力划分任务,并在比实际遇到的更干净、更稳定的环境中评估智能体。我们推出DuMateBench,这是一个从从大型生产智能体平台收集的匿名化且经过隐私筛查的用户会话重建的真实会话基准。每个任务都保留了相关的解决方案前交互历史、持久配置和工作空间状态,随后通过人工验证。最终的基准包含200个任务,涵盖8个广泛场景和17个细粒度能力类别,大多数任务需要多种能力协调。我们在注入了三种真实环境复杂性形式(不足、不稳定、有噪声)的隔离Docker容器中执行这些任务,并使用混合确定性与大语言模型(LLM)作为评判者的评估协议评估性能。对五个代表性自主智能体框架搭配四个最先进的LLM开展的实验显示,严格任务完成存在显著差距。补充的鲁棒性、效率和诊断分析进一步表明,环境扰动下的性能由LLM的能力和周围智能体框架共同塑造。代码和数据可在该httpsURL公开获取。

英文摘要

Autonomous agents are increasingly adopted to complete complex, multi-tool workflows in real-world settings. However, existing benchmarks typically separate tasks by application or capability and evaluate agents in environments that are cleaner and more stable than those encountered in practice. We introduce DuMateBench, a real-session benchmark reconstructed from anonymized and privacy-screened user sessions collected from a large-scale production agent platform. Each task preserves the relevant pre-solution interaction history, persistent configurations, and workspace state, and is then validated through human verification. The resulting benchmark comprises 200 tasks spanning 8 broad scenarios and 17 fine-grained capability categories, with most tasks requiring multiple capability coordination. We execute these tasks in isolated Docker containers injected with three forms of real-world environmental complexity: Insufficient, Unstable, and Noisy, and assess performance using a hybrid deterministic and LLM-as-Judge evaluation protocol. Experiments across five representative autonomous-agent frameworks paired with four state-of-the-art LLMs reveal substantial gaps in strict task completion. Complementary robustness, efficiency, and diagnostic analyses further show that performance under environmental perturbations is jointly shaped by the capabilities of the LLM and the surrounding agent framework. The code and data are publicly available at https://dumatebench.com/.

发表机构

  • Renmin University of China(中国人民大学)
  • Shandong University(山东大学)
  • Michigan State University(密歇根州立大学)
  • Nankai University(南开大学)
  • East China Normal University(华东师范大学)
  • Imperial College London(伦敦帝国学院)
  • Baidu, Inc.(百度公司)

机构由 AI 辅助整理,请以论文原文为准。

↑