arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Eon基准的Era:一个具有精确真实标签的生成式企业资产,用于基准测试LLM智能体

The Era by Eon Benchmark: A Generated Enterprise Estate with Exact Ground Truth for Benchmarking LLM Agents

Benjamin Gruenbaum, Doron Porat, Assaf Natanzon, Roy Zavida, Chen Dinachi, Or Itzahary

arXiv 2609.09853首次发表:更新:

发表机构

Eon(Eon)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该基准通过虚构企业资产和精确答案键,为评估企业工具LLM智能体提供真实标签,显著提升真实感得分。

AI 中文摘要

面向企业记录系统的LLM智能体无法在客户生产数据上进行评估,且现有替代方案均不提供真实标签。我们提出了Era by Eon基准,用于评估使用企业工具的LLM智能体。该基准围绕一个完整的虚构公司构建,包含产品模拟器、公司特定内部数据库、基准问题以及计算得出的答案键。行业、公司规模、商业模式、应用组合以及一个种子定义了每家公司。一个种子实体图向Salesforce、Zendesk、Slack、Gong及其他产品的模拟器提供共享公司数据。一个基于问题的生成器创建内部数据库的模式和记录,它先从同一图中获取共享实体、键和值,再生成数据库特定的事实。因此,这两种机制描述了一个一致的企业资产。每个预期答案均根据最终记录计算得出,因此评分是精确的。设计和答案键检查验证了内部数据库,而真实感评分卡和对抗性检测器验证了实体图。在23家生成的公司中,平均真实感得分从61.8升至97.0,且没有记录被标记为合成。在报告的模拟器轨道比较中,九个模型对相同的33个问题各回答三次,准确率估计范围从42.4%到76.8%,且在修正后,36个成对差异中仍有三个保持显著。

英文摘要

LLM agents for enterprise systems of record cannot be evaluated on customer production data, and no existing substitute provides ground truth. We present the Era by Eon Benchmark for evaluating LLM agents that use enterprise tools. The benchmark is built around a complete fictional company. It includes product simulators, company-specific internal databases, benchmark questions, and computed answer keys. Industry, company size, business model, application portfolio, and a seed define each company. One seeded entity graph supplies shared company data to simulators of Salesforce, Zendesk, Slack, Gong, and other products. A questionconditioned generator creates the schemas and records for internal databases. It takes shared entities, keys, and values from the same graph before generating database-specific facts. Both mechanisms therefore describe one consistent enterprise estate. Every expected answer is computed from the final records, so grading is exact. Design and answer-key checks validate the internal databases. A realism scorecard and adversarial detector validate the entity graph. Across 23 generated companies, the mean realism score rose from 61.8 to 97.0, with zero records flagged as synthetic. In the reported simulator-track comparison, nine models answered the same 33 questions three times each. Accuracy estimates ranged from 42.4% to 76.8%, and three of 36 pairwise differences remained supported after correction.

Comments12 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑