arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GDPevo:在真实商业任务上评估智能体的自进化能力

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks

Leijun Zhou, Zhihao Liu, Xiang Qu, Chenxu Liu, Yifei Liu, Yanke Yu, Jingzhe Xu, Xuejun Wu, Buyue Qian, Xi Chen, Yaowei Zheng, Junhao Hu

arXiv 2608.03764首次发表:更新:

AI 中文总结

本文提出GDPevo基准,通过规则杂交设计任务实现可归因评估,发现自进化可提升智能体测试准确率但仍远未达上限,相关资源已公开。

AI 中文摘要

智能体自进化是指智能体基于过往经验更新其持久状态,并复用该状态以更高效地解决相关任务。对自进化能力的评估存在诸多难点:现有基准对具有经济价值的任务领域覆盖有限,未始终将训练与测试任务设计为测试阶段的性能提升可归因于训练经验的形式,且仍易受数据污染影响。本文提出GDPevo,一个基于GDP相关企业工作流、原生支持进化的基准,以及生成该基准的全自动化数据流水线。其核心机制为规则杂交,即把每个企业工作流分解为原子级业务规则,将这些规则的子集分配至训练任务中,并在保留的测试任务中对其进行重组,从而使测试阶段的性能提升可归因于训练阶段的经验。GDPevo涵盖CRM、ERP、金融、医疗、法律及以数据为中心的工作流。其V1版本包含12个组共120个任务,每组含5个训练任务和5个保留测试任务。全自动化使该流水线可在两天内将任务套件扩展至24个组共240个任务(V2版本),为应对数据污染提供了可行方案。我们利用GDPevo评估了4个智能体,每个智能体由一个控制框架(harness)和一个模型组成,评估涵盖4种监督类型。结果显示,自进化使保留测试集的准确率持续提升,最高达16.44个百分点,但表现最佳的进化智能体仍远低于完全知情的基准上限91.6%,表明当前智能体的自进化能力仍远未得到充分发挥。我们在该httpsURL处公开发布此流水线、基准及完整评估结果。

英文摘要

Agent self-evolution updates an agent's persistent state from prior experience and reuses it to solve related tasks more effectively. Evaluating self-evolution is difficult: existing benchmarks provide limited coverage of economically valuable task domains, do not always design training and test tasks such that test-time gains can be attributed to training experience, and remain vulnerable to data contamination. We present GDPevo, an evolution-native benchmark grounded in GDP-related enterprise workflows, together with the fully automated data pipeline that generates it. Its core mechanism, rule hybridization, decomposes each enterprise workflow into atomic business rules, distributes subsets of these rules across training tasks, and recombines them in held-out test tasks so that test-time gains are attributable. GDPevo spans CRM, ERP, finance, healthcare, legal, and data-centric workflows. Its V1 release contains 120 tasks in 12 groups, with five training and five held-out test tasks per group. Full automation enables the pipeline to expand the suite to 240 tasks in 24 groups (V2) within two days, providing a practical response to contamination. Using GDPevo, we evaluate four agents, each comprising a harness and a model, under four supervision types. Self-evolution consistently improves held-out accuracy by up to 16.44 percentage points. But the best evolved agents remain far below the fully informed oracle ceiling of 91.6%, indicating that the self-evolution ability of current agents remains far from fully realized. We publicly release the pipeline, benchmark, and full evaluation results at https://github.com/Prism-Shadow/GDPevo.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑