发表机构
Duke University; Megagon Labs(杜克大学; Megagon 实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对文本到SQL中用户问题歧义导致错误的问题,提出结构化消歧新范式,构建首个含自然歧义的基准ARCS,实验显示现有模型准确率低,凸显该挑战。
AI 中文摘要
随着文本到SQL系统从演示阶段迈向实际部署,用户问题中的歧义成为错误的主要来源。这些歧义通常较为微妙,且与特定领域或数据相关,可能悄无声息地导致系统输出偏离用户的真实意图。传统上,歧义通过对话式澄清来解决,但这往往效率低下、认知负担重,且与现实用户工作流程契合度差。我们提出结构化消歧,这是一种新范式,通过明确且受限的交互而非自由形式的对话来解决歧义。我们构建了ARCS(SQL歧义消解语料库),这是首个针对真实世界数据库、包含自然发生且不受约束的歧义的文本到SQL基准,并完整标注了所有有效歧义点、解释及SQL查询。实验结果表明,在存在歧义的情况下,文本到SQL仍具挑战性:gpt-6-sol仅达到51%的端到端执行准确率,且没有任何开源模型超过27%。
英文摘要
As text-to-SQL systems move beyond demonstrations toward real-world deployment, ambiguity in user questions becomes a primary source of errors. Such ambiguities are often subtle, domain- or data-specific, and can silently cause system outputs to deviate from the user's true intent. Ambiguity is traditionally addressed through conversational clarification, which is often inefficient, cognitively demanding, and poorly aligned with real-world user workflows. We propose structured disambiguation, a new paradigm in which ambiguity is resolved through explicit, constrained interactions rather than free-form dialogue. We construct ARCS (Ambiguity Resolution Corpus for SQL), the first text-to-SQL benchmark featuring naturally occurring, unconstrained ambiguities over real-world databases, with complete annotations of all valid ambiguity points, interpretations, and SQL queries. Experimental results show that text-to-SQL remains challenging in the presence of ambiguity: gpt-6-sol achieves only 51% end-to-end execution accuracy, and no open-source model exceeds 27%.