arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ESF-Bench:针对现实世界企业应用的具有挑战性的槽填充场景基准测试

ESF-Bench: Benchmarking Challenging Slot-Filling Scenarios for Real-World Enterprise Applications

Toby Liang, Gopal Sarda, Sagar Davasam, Vikas Yadav

arXiv 2607.23326首次发表:更新:

发表机构

ServiceNow(ServiceNow公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对大语言模型在企业实际部署中槽填充面临的挑战,介绍ESF-Bench基准,它含多轮样本和槽,跨越多个领域,揭示了当前模型局限性,还公开了数据集、分类法及评估代码以推动该领域研究。

AI 中文摘要

大语言模型在企业中的应用迅速兴起,但在实际部署中因复杂系统约束和意外用户行为面临独特挑战。槽填充对于将非结构化输入转化为结构化、可操作数据至关重要。本文介绍了ESF-Bench,一个具有挑战性的企业槽填充基准,包含810个多轮样本和6530个槽,跨越8个独特领域。该基准揭示了当前大语言模型的显著局限性,如GPT-OSS-120b仅能成功提取20.7%基准样本的槽。为支持该领域研究,相关数据集、分类法及评估代码已在GitHub上公开发布。

英文摘要

The rapid rise of large language models (LLMs) has driven transformative adoption across enterprises. However, deploying these models in real-world settings presents unique challenges due to complex system constraints and unexpected user behaviors. Among these applications, slot filling is essential for converting unstructured input into structured, actionable data. In this work, we introduce ESF-Bench, a challenging Enterprise Slot Filling benchmark consisting of 810 multi-turn samples and 6530 slots over 8 unique domains. Curated using a taxonomy of the 57 most challenging slot-filling scenarios observed during real-world enterprise deployments, ESF-Bench exposes notable limitations in current state-of-the-art LLMs, with GPT-OSS-120b low successfully extracting slots for only 20.7% of benchmark samples. To support continued research in this area, we publicly release the benchmark dataset, taxonomy, and accompanying evaluation code on GitHub.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑