发表机构
Shandong University; Zhongguancun Academy; Beijing Institute of Technology; Tsinghua University(山东大学; 中关村学院; 北京理工大学; 清华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
EnterpriseBench提出一个企业级基准,涵盖静态问答到动态决策,含咨询、啤酒游戏和企业数字孪生三个交互场景,实验显示现有LLM智能体在企业任务中尚缺乏稳定可靠性。
AI 中文摘要
LLM智能体日益被期望支持企业工作流程,这些任务通常涉及信息缺失、不确定性、反馈和长期权衡。然而,现有的企业和金融基准主要测试静态能力,如信息提取、数值计算、领域知识和金融问答,而交互式和长期决策尚未得到充分探索。为弥补这一差距,我们引入了EnterpriseBench,一个评估LLM智能体从静态问答到动态决策全谱系能力的基准。具体而言,EnterpriseBench将现有的企业和金融问答数据集重组为一个统一的基础套件,按能力和难度进行标注,并引入了三个专业交互设置:咨询(基于管理咨询风格的商业案例,通过多轮信息寻求进行客户问题诊断);啤酒游戏(改编自经典供应链管理模拟,用于延迟反馈下的库存控制);以及企业数字孪生(一个基于项目的商业模拟器,用于劳动力、风险和项目规划)。在四个骨干模型下使用九种智能体方法进行的实验表明,当前智能体在企业场景中尚未实现稳定、全面和跨任务可靠性。这些结果表明,EnterpriseBench为在现实企业战略推理和决策中评估LLM智能体提供了一个实用基准。
英文摘要
LLM agents are increasingly expected to support enterprise workflows, where tasks often involve missing information, uncertainty, feedback, and long-term trade-offs. However, existing enterprise and financial benchmarks mainly test static capabilities such as information extraction, numerical calculation, domain knowledge, and financial QA, leaving interactive and long-horizon decision-making underexplored. To bridge this gap, we introduce EnterpriseBench, a benchmark that evaluates LLM agents across this spectrum, from static question answering to dynamic decision-making. Specifically, EnterpriseBench reorganizes existing enterprise and financial QA datasets into a unified foundational suite annotated by capability and difficulty, and introduces three professional interactive settings: Consulting, based on management-consulting-style business cases for client problem diagnosis through multi-turn information seeking; the Beer Game, adapted from a classic supply-chain management simulation for inventory control under delayed feedback; and Enterprise Digital Twin, a project-based business simulator for workforce, risk, and project planning. Experiments with nine agent methods under four backbone models show that current agents have not yet achieved stable, comprehensive, and cross-task reliability in enterprise scenarios. These results show that EnterpriseBench provides a practical benchmark for evaluating LLM agents in realistic enterprise strategic reasoning and decision-making.