重新审视Agentic-SQL:面向LLM文本转SQL的基于自主性的分类与实证基准分析
Agentic-SQL Revisited: Autonomy-Based Taxonomy and Empirical Benchmark Analysis for LLM Text-to-SQL
浏览论文内容
中文总结 AI 辅助
本研究针对LLM文本转SQL领域,构建基于自主性的分类框架,通过Spider案例研究分析8B开源模型与DeepSeek-V3等基线的表现,发布工具链用于构建可追溯的基准排行榜。
中文摘要 AI 辅助
基于大语言模型(LLM)的文本转SQL技术在异构基准、主干模型及推理协议上取得进展,导致跨系统比较脆弱。我们将该领域重构为排行榜聚合:收集作者自行报告的指标,沿推理自主性轴组织,涵盖受限、上下文内、迭代、Agentic及推理内化生成,每个条目都有可追溯来源。为实证锚定该聚合,我们针对Spider开展聚焦案例研究,对比8B开源主干模型(有无思维链(CoT)监督)与少样本DeepSeek-V3、GLM-4基线。得出四个模式:Spider向BIRD和Spider-2.0的迁移不均衡;自主性以显著成本换取鲁棒性;推理内化介于仅答案解码与外部协调Agent之间;CoT增益集中于困难及极困难查询。我们发布Python工具链,镜像自主性轴,以便未来方法可直接加入排行榜。
英文摘要
LLM-based Text-to-SQL progress is reported across heterogeneous benchmarks, backbones, and inference protocols, making cross-system comparison fragile. We reframe the field as a leaderboard aggregation: we collect the metrics authors themselves report and organize them along an inference-autonomy axis spanning constrained, in-context, iterative, agentic, and reasoning-internalized generation, with traceable provenance for every cell. To anchor the aggregation empirically, we run a focused case study on Spider, comparing 8B open-source backbones with and without chain-of-thought (CoT) supervision against few-shot DeepSeek~V3 and GLM-4 baselines. Four patterns emerge: Spider gains transfer unevenly to BIRD and Spider~2.0; autonomy buys robustness at non-trivial cost; reasoning internalization sits between answer-only decoding and externally orchestrated agents; and CoT gains concentrate on Hard and Extra-Hard queries. We release a Python harness mirroring the autonomy axis so that future methods can be added directly to the leaderboard.