在链歧义与意图漂移场景下评估大型语言模型(LLMs)的对话式文本转SQL能力
Evaluating LLMs on Conversational Text-to-SQL under Chain Ambiguity and Intent Drift
- The Hong Kong Polytechnic University(香港理工大学)
- City University of Macau(澳门城市大学)
- Jilin University(吉林大学)
- Beihang University(北京航空航天大学)
- Jinan University(暨南大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究针对对话式文本转SQL中用户意图演变未被现有基准覆盖的问题,推出TIDE-Bench基准并评估12个LLMs,发现链识别瓶颈等问题,代码已发布。
AI中文摘要:
大型语言模型(LLMs)的近期进展已使对话式文本转SQL成为用户与数据库间的实用接口,通常涉及多轮澄清与修正。然而现有基准主要评估执行准确率,未覆盖多轮对话中用户意图的演变与转变。为解决此问题,我们推出TIDE-Bench,一个针对链歧义与意图漂移场景的对话式文本转SQL评估基准,聚焦两类重复出现的模式:链歧义(欠规范问题触发带条件依赖的分层澄清)与意图漂移(用户撤回并替换先前承诺的请求元素)。TIDE-Bench基于BIRD的514个锚定SQL构建,包含1542个样本,除执行准确率外还引入链识别、漂移识别-解析的专用指标。对12个先进LLMs的评估显示,存在不受澄清频率影响的持续链识别瓶颈、较大的漂移识别-解析差距,以及两种失效模式同时激活时的重叠情况。TIDE-Bench的对应代码已发布以供进一步研究。
英文摘要:
Recent advances in large language models (LLMs) have established conversational text-to-SQL as a practical interface between users and databases, often involving multiple turns of clarification and revision. However, existing benchmarks primarily evaluate execution accuracy, leaving the unfolding and shifting of user intent across turns largely uncovered. To address this, we introduce TIDE-Bench, a benchmark for conversational text-to-SQL under chain ambiguity and intent drift evaluation, targeting two recurring patterns: chain ambiguity, where an underspecified question triggers layered clarification with conditional dependencies, and intent drift, where the user retracts and replaces a previously committed request element. Built on 514 anchor SQLs from BIRD, TIDE-Bench comprises 1,542 samples and introduces dedicated metrics for chain identification and drift recognition-resolution beyond execution accuracy. Evaluating 12 advanced LLMs reveals a persistent chain identification bottleneck unaffected by clarification frequency, a wide drift recognition-resolution gap, and overlap between failure modes when jointly activated. The corresponding code of TIDE-Bench is released for further research.