arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

这些模块是否物有所值?对上下文学习文本到SQL的范式级准确性-成本分析

Are These Modules Worth Their Cost? A Paradigm-Level Accuracy-Cost Analysis of In-context Learning Text-to-SQL

Jiayan Lin, Yujia Liu, Zijin Hong, Zheng Yuan, Yilin Xiao, Hao Chen, Qinggang Zhang, Xiao Huang, Feiran Huang

arXiv 2608.28432首次发表:更新:

发表机构

Jinan University; The Hong Kong Polytechnic University; City University of Macau; Jilin University; Beihang University(暨南大学; 香港理工大学; 澳门城市大学; 吉林大学; 北京航空航天大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对上下文学习文本到SQL流程的模块,量化各范式的成本-准确性贡献,发现执行反馈优化性价比最高,提出感知成本的分层指南并可迁移至其他主干模型。

AI 中文摘要

上下文学习(ICL)文本到SQL的最新进展,通过围绕基础生成器构建日益复杂的流程,大幅提升了公共基准上的执行准确率,但现有研究通常仅报告聚合的端到端准确率,未量化各个设计选择的边际准确性-成本贡献。因此,提供统一的范式级成本-准确性量化仍是理解和配置现代文本到SQL的关键挑战。为解决该问题,我们在单一受控实现下,针对ICL文本到SQL流程的五个重复模块实例化17种范式级配置,并在四种覆盖不同能力水平和推理风格的主干模型上,归因每种范式的边际贡献及产生的成本。我们的分析显示,执行反馈优化是唯一在成本持续较低时仍能普遍发挥作用的范式,而大多数其他模块仅在依赖主干模型的条件下才有用。Token计数表明,输入需求与流程结构的关联更紧密,而输出需求对主干模型的生成行为更敏感。跨模块分析进一步显示,堆叠(stacking)可提升大多数主干模型的准确率,不过增益的组合方式随主干模型能力变化。我们还发现,固定预算通常更适合用于在中端主干模型上构建更复杂的流程,而非升级为使用精简流程的前沿模型。这些发现提炼为可操作的、感知成本的分层指南,可迁移到另外五种主干模型,无需进行逐范式搜索。

英文摘要

Recent advances in in-context learning (ICL) text-to-SQL have substantially improved execution accuracy on public benchmarks by assembling increasingly elaborate pipelines around the base generator, yet existing studies typically report aggregate end-to-end accuracy, without quantifying the marginal accuracy-cost contribution of individual design choices. Consequently, providing a unified, paradigm-level cost-accuracy quantification remains a critical challenge for understanding and configuring modern text-to-SQL. To address this, we instantiate 17 paradigm-level configurations across five recurring modules of the ICL text-to-SQL pipeline under a single controlled implementation, and attribute each paradigm's marginal contribution and incurred cost across all four backbones spanning diverse capability levels and reasoning styles. Our analysis reveals that execution-feedback refinement is the only paradigm whose benefit holds universally at consistently low cost, while most other modules help only under backbone-dependent conditions. Token accounting shows that input demand is more closely tied to pipeline structure, whereas output demand is more sensitive to backbone generation behavior. Cross-module analysis further shows that stacking improves accuracy on most backbones, although how the gains compose varies with backbone capability. We also find that a fixed budget is often better spent engineering a more elaborate pipeline over a mid-tier backbone than upgrading to a frontier model with a lean pipeline. These findings distill into an actionable, cost-aware tiered guideline that transfers to five additional backbones without per-paradigm search.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑