arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31906cs.AI

EmailBench:用于评估LLM智能体在企业电子邮件和生产力任务上的基准

EmailBench: A Benchmark for Evaluating LLM Agents on Enterprise Email and Productivity Tasks

Mukul Singh, Mansi Uniyal, Devin Devlin, Wen Xie, Big Thadawasin, Ritam Dutt, Vivian Lai, Hyeonsu B. Kang

首次发表
浏览论文内容

中文总结 AI 辅助

EmailBench是一个包含206个场景、16个任务类别的基准,用于评估LLM智能体在企业电子邮件任务中的表现,发现工具调用成功不等于任务完成,最佳配置仅通过33.5%的场景。

中文摘要 AI 辅助

企业电子邮件智能体必须结合信息检索、结构化状态变更、时间推理和多步骤协调。最近的智能体基准包括生产力任务,但很少有在自包含环境中以类型化电子邮件工作流为中心的基准。我们引入了EmailBench,一个包含16个任务类别、206个电子邮件和生产力场景的基准。该基准将类型化电子邮件API规范与提供商无关的命名、一个确定性的合成Enron风格语料库以及一个场景套件相结合,其主题选择参考了来自交互式原型的聚合任务意图遥测。其混合评估协议结合了258个可执行静态断言和211个LLM评分标准。我们在一个固定的单用户语料库上评估了八个LM配置。最佳配置仅通过了33.5%的场景,尽管其99.7%的工具调用在未观察到API失败的情况下完成,且通过率在不同任务类别间差异显著。这一差距表明,有效的工具执行并不等同于任务完成。EmailBench为端到端电子邮件智能体评估提供了一个自包含环境,更广泛的工具覆盖、多角色测试和重复运行评估是未来的工作方向。

英文摘要

Enterprise email agents must combine information retrieval, structured state changes, temporal reasoning, and multi-step coordination. Recent agent benchmarks include productivity tasks, but few center on typed email workflows in a self-contained environment. We introduce EmailBench, a benchmark of 206 email and productivity scenarios across 16 task categories. The benchmark couples a typed email API specification with provider-neutral naming, a deterministic synthetic Enron-inspired corpus, and a scenario suite whose topic selection was informed by aggregate task-intent telemetry from an interactive prototype. Its hybrid evaluation protocol combines 258 executable static assertions with 211 LLM rubrics. We evaluate eight LM configurations on a fixed single-user corpus. The best-performing configuration passes only 33.5% of scenarios despite 99.7% of its tool calls completing without an observed API failure, with pass rates varying substantially across task categories. This gap shows that valid tool execution is not equivalent to task completion. EmailBench provides a self-contained environment for end-to-end email-agent evaluation, with broader tool coverage, multi-persona testing, and repeated-run evaluation as future work areas.

补充信息

↑