AI 中文总结
研究针对文本到SQL系统在处理现实模糊问题的不足,提出统一分类法、多代理生成管道及动态模拟环境ABISS。通过实验揭示模型在子类别分类和澄清条件下SQL生成的瓶颈,提供真实类别虽有收益,但模糊问题执行率仍低,且发布相关代码。
AI 中文摘要
大语言模型在精心策划的文本到SQL基准测试中表现出高性能;然而,现实世界中的用户经常提出模糊或无法回答的问题,当前系统处理得很差。三个相互关联的差距阻碍了进展:不完整的分类法、针对现实世界场景的现实基准测试生成以及静态用户交互。我们通过三项贡献解决了上述所有问题:(1)一个统一的分类法,涵盖8个类别,包括模糊和无法回答的问题;(2)一个多代理生成管道,具有两阶段过程(自然语言问题生成,然后是SQL基础)和一个明确的类别一致性验证阶段,从由本地开源模型委员会验证的任意数据库中生成问题;(3)ABISS(使用交互模拟会话的模糊基准测试),一个动态模拟环境,文本到SQL代理在多轮对话中与风格感知的模拟用户进行交互。在ABISS-BIRD和ABISS-Spider上对八个开源模型进行的实验揭示了两个基本瓶颈。第一个是子类别分类:模型检测到问题有问题,但难以确定具体的子类别。第二个是澄清条件下的SQL生成:即使在收到有用的用户信息后,模型在最终解决步骤中也经常失败。提供真实类别在两个数据集的执行和反馈方面都有很大收益,但即使在先知类别标签下,模糊问题的执行率仍然很低。我们在GitHub上发布了用于数据生成和基准测试的代码(此https URL)。
英文摘要
Large Language Models (LLMs) demonstrate high performance on curated Text-to-SQL benchmarks; nevertheless, real-world users frequently pose ambiguous or unanswerable questions that current systems handle poorly. Three interconnected gaps hinder progress: incomplete taxonomies, realistic benchmark generation for real-world settings, and static user interaction. We address all of the above issues through three contributions: (1) a unified taxonomy of 8 categories covering ambiguous and unanswerable questions; (2) a multi-agent generation pipeline with a two-stage process (NLQ generation followed by SQL grounding) and an explicit Category Conformance validation stage, producing questions from arbitrary databases validated by a council of local open-source models; and (3) ABISS (Ambiguity Benchmark using Interaction-Simulated Sessions), a dynamic simulation environment where Text-to-SQL agents interact with style-aware simulated users across multi-turn dialogues. Experiments with eight open-source models on ABISS-BIRD and ABISS-Spider reveal two fundamental bottlenecks. The first is subcategory classification: models detect that a question is problematic yet struggle to pinpoint the specific subcategory. The second is clarification-conditioned SQL generation: even after receiving useful user information, models often still fail in the final resolution step. Providing the ground truth category yields large gains in both execution and feedback across both datasets, yet ambiguous-question execution remains low even under oracle category labels. We release our code for data generation and benchmark on GitHub (https://github.com/giosullutrone/ABISS-Evaluating-Text-to-SQL-Systems-Through-Agent-Interaction).