arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.29345cs.AIcs.CL

BIRD-History:带有细粒度知识标注的历史驱动型Text-to-SQL基准测试

BIRD-History: A Benchmark for History-Driven Text-to-SQL with Fine-Grained Knowledge Annotations

发表机构浙江大学
查看机构详情
  • Zhejiang University(浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

Yunfan Zhou, Qiming Shi, Yizhou Yang, Di Weng, Yingcai Wu

首次发表
浏览论文内容

中文总结 AI 辅助

本研究推出带细粒度知识标注的BIRD-History基准测试,针对现有Text-to-SQL系统无法利用历史查询日志处理隐含领域知识的问题,提出插件式检索器,可提升多个Text-to-SQL系统的未明确查询处理性能。

中文摘要 AI 辅助

尽管近期基于大型语言模型(LLM)的Text-to-SQL系统在标准基准测试上取得了出色性能,但当用户查询隐含依赖于特定领域知识时,它们仍表现不佳,这类知识包括业务逻辑、数据约定和分析实践,既未被模式捕获,也未在自然语言问题中明确表述。历史SQL查询日志是这类知识的宝贵来源,但现有基准测试无法充分支持对历史驱动方法的评估。为解决这一缺口,我们推出BIRD-History,该基准测试包含11个数据库中的1393个任务,旨在评估Text-to-SQL系统利用历史SQL脚本为未明确的自然语言问题提供依据的能力。每个任务都标注了真实标签,指定哪些历史查询包含相关知识,以及哪些SQL子句对其进行编码,从而支持对检索有效性和知识利用的系统评估。除基准测试外,我们还提出一种插件式检索器,该检索器从历史SQL脚本中提取五类外部知识,随后检索并重新排序相关片段以生成查询。该检索器可无缝集成到现有少样本Text-to-SQL流程中,无需修改提示词。实验表明,其在四个Text-to-SQL系统上均实现了一致的性能提升,凸显了利用历史查询日志处理未明确查询的价值。数据集和代码已在该URL开源。

英文摘要

While recent Large Language Model (LLM)-based text-to-SQL systems achieve impressive performance on standard benchmarks, they struggle when user queries implicitly rely on domain-specific knowledge, such as business logic, data conventions, and analytical practices, that is neither captured by the schema nor explicitly stated in the natural language question. Historical SQL query logs offer a valuable source of such knowledge, yet existing benchmarks do not adequately support evaluation of history-driven approaches. To address this gap, we introduce BIRD-History, a benchmark consisting of 1,393 tasks across 11 databases, designed to evaluate text-to-SQL systems' ability to ground underspecified natural language questions using historical SQL scripts. Each task is annotated with ground-truth labels specifying which historical queries contain relevant knowledge and which SQL clauses encode it, enabling systematic evaluation of both retrieval effectiveness and knowledge utilization. Alongside the benchmark, we propose a plug-in retriever that extracts five types of external knowledge from historical SQL scripts, then retrieves and reranks relevant fragments for query generation. The retriever integrates seamlessly into existing few-shot text-to-SQL pipelines without requiring prompt modifications. Experiments demonstrate consistent improvements across four text-to-SQL systems, highlighting the value of leveraging historical query logs for handling underspecified queries. Dataset and code are open-sourced on https://github.com/zjuidg/BIRD-History.

补充信息

↑