ERPBench:评估大型语言模型智能体在竞争市场生态下的企业决策能力
ERPBench: Evaluating LLM Agents for Enterprise Decision-Making Across Competitive Market Ecologies
浏览论文内容
中文总结 AI 辅助
本研究推出带执行检测的ERPBench基准,评估LLM智能体在两种竞争市场生态的企业决策能力,发现模型排名在生态间迁移性差,为企业智能体评估提供新基准。
中文摘要 AI 辅助
大型语言模型(LLM)智能体越来越多地被提出用于企业工作流程,但现有评估很少测试商业决策结论是否能在竞争市场生态中迁移。我们推出ERPBench,这是一个带执行检测的企业决策智能体基准,基于六轮企业资源规划(ERP)模拟,涵盖耦合定价、生产、采购、库存、财务及共享市场竞争。ERPBench在两个匹配的竞争市场生态中评估相同的100个固定问题:Solo生态中,每个被评估的LLM智能体与固定基于规则的对手竞争;Arena生态中,六个被评估的LLM智能体在共享市场中竞争。涉及六个模型系列,共产生1200个模型级轨迹,涵盖7200个决策轮次。在观察到的服务配置下,领先模型因生态而异:DeepSeek在Solo中领先(平均估值2.5229亿,平均排名1.67),而Gemini在Arena中领先(平均估值2.6395亿,平均排名1.76)。两个生态仅在100个问题中的21个上识别出相同的任务级胜者,Gemini的垫底率在Arena中从22%降至0%。ERPBench支持配对评估企业智能体排名是否能在竞争市场生态中迁移,并辅以聚合执行干预分析。代码和基准资源可在我们的this https URL获取。
英文摘要
Large language model (LLM) agents are increasingly proposed for enterprise workflows, yet existing evaluations rarely test whether business-decision conclusions transfer across competitive market ecologies. We introduce ERPBench, an execution-instrumented benchmark for enterprise decision agents in a six-round Enterprise Resource Planning (ERP) simulation with coupled pricing, production, procurement, inventory, finance, and shared-market competition. ERPBench evaluates the same 100 fixed problems in two matched competitive market ecologies: Solo, where each evaluated LLM agent competes against fixed rule-based opponents, and Arena, where six evaluated LLM agents compete in a shared market. Across six model families, this yields 1,200 model-level trajectories spanning 7,200 decision rounds. Under the observed service configuration, the leading model differs between ecologies: DeepSeek leads in Solo (252.29M mean valuation; mean rank 1.67), whereas Gemini leads in Arena (263.95M; 1.76). The two ecologies identify the same task-level winner on only 21 of 100 problems, and Gemini's bottom-rank rate falls from 22 % to 0 % in Arena. ERPBench supports paired evaluation of whether enterprise-agent rankings transfer across competitive market ecologies, supplemented by aggregate execution-intervention analysis. Code and benchmark resources are available in our https://github.com/GAIR-NLP/erp-bench.
发表机构
- Shanghai Innovation Institute(上海创新研究院)
- Beijing Institute of Technology(北京理工大学)
- Shanghai Jiao Tong University(上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。