ERPBench:企业软件中计算机使用智能体的状态基础评估范式
ERPBench: A State-Grounded Evaluation Paradigm for Computer-Use Agents in Enterprise Software
- Accenture(埃森哲)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对ERP系统对计算机使用智能体的独特挑战,提出ERPBench基准,在实时ERP系统上按数据库真实值评分,发现强通用GUI性能不转化为企业可靠性,并刻画特有失败模式。
AI中文摘要:
通过截图和模拟操作运行的计算机使用智能体正在快速发展,然而对其评估仍主要局限于通用桌面和网页任务。企业资源规划(ERP)系统支撑着全球组织的财务、采购、库存和客户运营,为计算机使用智能体带来了独特挑战:密集的界面、协调的多步骤交互,以及错误会改变持久业务记录而非在屏幕上显现。现有的企业基准依赖专有平台或此类软件的模拟近似。我们引入了ERPBench,一个在实时且可复现的ERP系统上评估仅基于截图的智能体的基准,并根据数据库中的真实值对每个任务进行评分。除基准外,我们提出了一个生产级测试平台,将智能体操作置于人工审批之后以确保安全部署,ERPBench可自主运行该平台。通过评估六个闭源和开源智能体,我们证明强大的通用图形用户界面性能并不能转化为企业可靠性。即使智能体到达正确的表单并保存,存储的记录也常常是错误的:一些智能体在高达85%的运行中保存了记录,但写入正确值的比例低至3%。我们进一步刻画了企业工作流特有的失败模式。
英文摘要:
Computer-use agents that operate through screenshots and simulated actions are advancing rapidly, yet their evaluation remains anchored to general desktop and web tasks. Enterprise Resource Planning systems run the finance, procurement, inventory, and customer operations of organizations worldwide, and pose distinct challenges for computer-use agents: dense interfaces, coordinated multi-step interactions, and errors that alter persistent business records rather than surfacing on screen. Existing enterprise computer-use benchmarks rely on proprietary platforms or on simulated approximations of such software. We introduce ERPBench, a benchmark that evaluates screenshot-only agents on a live and reproducible system and scores each task against ground-truth values in its database. Beyond the benchmark, we present a production-grade harness that gates agent actions behind human approval for safe deployment. Evaluating six closed and open-source agents, we demonstrate that strong general performance does not transfer to enterprise reliability. Even when an agent reaches the right form and saves it, the stored record is often wrong: some agents save in up to 85% of runs but write the correct value in as few as 3%. We further characterize failure modes specific to enterprise workflows.