BC-Bench:在ERP领域特定语言中评估智能体工程
BC-Bench: Evaluating Agentic Engineering in a Domain-Specific Language for ERP
- Microsoft(微软公司)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究推出BC-Bench基准,评估智能体工程在ERP领域DSL(AL)上的表现,发现通用基准的改进未一致迁移到AL,凸显领域特定评估的必要性。
AI中文摘要:
智能体工程系统在通用基准测试中表现出强大性能,但其在企业资源规划(ERP)领域特定语言(DSL)中的有效性仍未得到充分探索。我们推出BC-Bench,这是一个用于评估在Microsoft Dynamics 365 Business Central的DSL(AL)上进行智能体工程的基准测试。BC-Bench包含从两个微软自有生产仓库中提取的101个手动整理的任务,反映了真实的ERP开发工作流程。我们采用SWE-Bench方法,解决AL生态系统的独特约束,包括有限的公共资源和复杂的环境配置。除了生成功能代码外,BC-Bench还评估测试生成,并支持通常包含视觉上下文的多模态问题陈述。我们在两个智能体框架上评估多个前沿模型,利用多运行指标考虑不确定性。在错误修复类别中,在我们评估的设置下,模型间解决率的差异大于两个评估智能体框架之间的差异,并且在通用基准测试中报告的改进并未一致地转移到AL。这些结果凸显了对领域特定评估的需求。
英文摘要:
Agentic engineering systems have shown strong performance on general-purpose benchmarks, yet their effectiveness in enterprise resource planning (ERP) domain-specific languages (DSLs) remains underexplored. We introduce BC-Bench, a benchmark designed to evaluate agentic engineering on real-world tasks in AL, the DSL for Microsoft Dynamics 365 Business Central. BC-Bench comprises 101 manually curated tasks extracted from two Microsoft-owned production repositories, reflecting authentic ERP development workflows. Adapting the SWE-Bench methodology, we address the unique constraints of the AL ecosystem---including limited public resources and complex environment provisioning. Beyond generating functional code, BC-Bench evaluates test generation and supports multimodal problem statements where visual context is commonly present. We evaluate multiple frontier models across two agent harnesses, utilizing multi-run metrics to account for nondeterminism. In the Bug Fixing category, under our evaluated settings, between-model differences in resolution rate are larger than differences between the two evaluated agent harnesses, and improvements reported on general-purpose benchmarks do not consistently transfer to AL. These results highlight the need for domain-specific evaluation.