发表机构
ServiceNow(ServiceNow公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对大语言模型在企业实际部署中槽填充面临的挑战,介绍ESF-Bench基准,它含多轮样本和槽,跨越多个领域,揭示了当前模型局限性,还公开了数据集、分类法及评估代码以推动该领域研究。
AI 中文摘要
大语言模型在企业中的应用迅速兴起,但在实际部署中因复杂系统约束和意外用户行为面临独特挑战。槽填充对于将非结构化输入转化为结构化、可操作数据至关重要。本文介绍了ESF-Bench,一个具有挑战性的企业槽填充基准,包含810个多轮样本和6530个槽,跨越8个独特领域。该基准揭示了当前大语言模型的显著局限性,如GPT-OSS-120b仅能成功提取20.7%基准样本的槽。为支持该领域研究,相关数据集、分类法及评估代码已在GitHub上公开发布。
英文摘要
The rapid rise of large language models (LLMs) has driven transformative adoption across enterprises. However, deploying these models in real-world settings presents unique challenges due to complex system constraints and unexpected user behaviors. Among these applications, slot filling is essential for converting unstructured input into structured, actionable data. In this work, we introduce ESF-Bench, a challenging Enterprise Slot Filling benchmark consisting of 810 multi-turn samples and 6530 slots over 8 unique domains. Curated using a taxonomy of the 57 most challenging slot-filling scenarios observed during real-world enterprise deployments, ESF-Bench exposes notable limitations in current state-of-the-art LLMs, with GPT-OSS-120b low successfully extracting slots for only 20.7% of benchmark samples. To support continued research in this area, we publicly release the benchmark dataset, taxonomy, and accompanying evaluation code on GitHub.