Era by Eon:在隐藏知识上对企业智能体进行基准测试
Era by Eon: Benchmarking Enterprise Agents on Hidden Knowledge
浏览论文内容
中文总结 AI 辅助
该研究在Era by Eon基准上新增隐藏事实问题,评估12个企业智能体,发现最佳智能体答对18/24,而最难问题所有智能体仅1/84成功,凸显其局限。
中文摘要 AI 辅助
在Era by Eon基准测试中,每个问题都陈述了其答案的规则,代码根据生成的公司数据计算答案。当智能体能够运行代码时,四个最强的模型各自在27个此类问题中答对22至25个,因此该基准测试几乎无法区分它们。我们添加了八个依赖隐藏事实的问题模板。没有任何问题或文档陈述隐藏事实,而看似包含该事实的记录显示的是其他内容。其他数据隐含了该事实。例如,销售系统称客户因时机问题取消了购买。在录音通话中,客户将责任归咎于一次中断。对于每个生成的公司,代码填充每个模板并计算精确答案,无需语言模型。我们评估了12个智能体。每个智能体将一个模型与一个智能体程序配对,该程序将模型连接到公司的系统。最佳智能体在其24次尝试(每个问题三次)中答对18次。六个模型中有四个在使用任何程序时最多答对24次中的6次。最难的问题需要从几个相似记录中挑选一个,例如客户签署了三个续约报价中的哪一个。所有智能体合计仅在84次尝试中答对了两个此类问题中的一次。
英文摘要
In the Era by Eon benchmark, each question states the rules for its answer, and code computes the answer from a generated company's data. When agents can run code, the four strongest models each answer 22 to 25 of 27 such questions, so the benchmark barely separates them. We add eight question templates that depend on hidden facts. No question or document states a hidden fact, and the records that seem to hold it show something else. Other data implies it. For example, the sales system says a customer dropped a purchase because of timing. On a recorded call, the customer blames an outage. For each generated company, code fills each template and computes an exact answer without a language model. We evaluate 12 agents. Each pairs a model with an agent program, which connects it to the company's systems. The best agent answers 18 of its 24 attempts, three per question, correctly. Four of the six models answer at most 6 of 24 with any program. The hardest questions require picking one of several similar records, such as which of three renewal offers a customer signed. All agents together answered two such questions correctly in only 1 of 84 attempts.