束搜索、自一致性以及小语言模型中语法约束下文本到SQL任务推理时间缩放的局限性
Beam Search, Self-Consistency, and the Limits of Inference-Time Scaling for Grammar-Constrained Text-to-SQL in Small Language Models
浏览论文内容
中文总结 AI 辅助
本文针对语法约束下的小语言模型文本到SQL任务,研究了模型规模与推理计算的权衡,发现受约束权衡表现不同,束搜索优于采样+投票,且增大模型规模比增加推理计算更有利。
中文摘要 AI 辅助
大语言模型使用中常见的权衡是减小模型规模,同时增加推理时的计算量,例如通过使用更宽的束搜索。本文研究了这种“模型规模与推理计算”权衡的约束情况,即模型输出在推理时受严格语法约束。我们的结果表明,受约束的权衡表现与无约束的权衡不同。我们研究将自然语言查询转换为等价SQL查询的任务(文本到SQL),在Spider文本到SQL基准上评估性能,使用Qwen2.5-Instruct模型系列,规模从0.5B到7B参数,均为4位精度。我们实验了两种改变推理计算的方法:(i)可变束数的束搜索;(ii)采样+投票,即采样多个受约束的输出,然后对其执行结果投票,采样数量可变。在包含1034个样本的开发集上,我们发现:(a)束搜索和采样+投票均能提高准确率,尤其在较小模型规模下;(b)本实验中“模型规模与推理计算”的权衡并无优势,因为在相同模型规模下,增大模型规模通常比增加推理计算带来更高的准确率;(c)在相同推理预算下,束搜索的表现优于采样+投票,这一结果尤为受关注,因为它与无约束权衡的发现形成对比。
英文摘要
One common trade-off in the use of large language models involves reducing the size of the model while increasing the amount of computation at inference time, for example by using a wider beam search. In this paper, we examine the constrained case of this "model size vs. inference compute" trade-off, in which the model outputs are constrained by a strict grammar at inference time. Our results demonstrate that the constrained trade-off behaves differently from the unconstrained trade-off. We investigate the task of converting a prose query into an equivalent SQL query (text-to-SQL). Performance is evaluated on the Spider text-to-SQL benchmark, using the Qwen2.5-Instruct model family ranging in size from 0.5B to 7B parameters, all at 4-bit precision. We experiment with two approaches to varying inference compute: (i) beam search with a variable number of beams; and (ii) sample+vote, i.e., sampling several constrained outputs and then voting on their execution results, where the number of samples is varied. On the 1034-example development set, we find that: (a) both beam search and sample+vote improve accuracy, especially on smaller model sizes; (b) the "model size vs.\ inference compute" trade-off is not advantageous in this experiment, because moving to a larger model size typically results in higher accuracy than increasing inference compute on the same model size; (c) beam search outperforms sample+vote at a matched inference budget. This latter result is of particular interest since it contrasts with the findings of the unconstrained trade-off.
发表机构
- Deep Network Understanding Lab(深度网络理解实验室)
- Dickinson College(狄金森学院)
机构由 AI 辅助整理,请以论文原文为准。