发表机构
Megagon Labs(Megagon Labs)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出BudgetSchemaBench,一个基于执行结果的诊断基准,用于评估文本到SQL中不同模式上下文预算和序列化方式对执行准确率的影响,发现预算比序列化方式更关键。
AI 中文摘要
基于结构化源的数据代理必须将数据库模式适配到模型的上下文窗口中。大型目录可能跨越多个数据库和数千列,因此成本约束可能要求在上下文窗口尚未填满之前,就在表覆盖范围与序列化细节之间做出选择。我们引入了BudgetSchemaBench,一个针对此设置的、基于执行结果的诊断基准。其构建过程从金标准SQL中机械地推导出相关性标签,无需人工或LLM撰写的真实标签。使用一个汇集了80个数据库的目录,我们扫描了四个模式上下文预算,并在保持每个检索器的表排名固定的情况下比较了三种表示方式。一个源命名空间检查会拒绝那些从错误数据库获得正确结果的查询。评估涵盖三种条件:端到端检索;冻结金标准(其中所需的表得到保证);以及一个移除这些表的探针。对于使用原始序列化的主要求解器,将预算从目录的2.5%提高到50%,在1,279个保留问题上,词法检索下的执行准确率提高了18个百分点,而密集检索下仅提高了3个百分点;密集检索器在最小预算下已经找到了大多数所需的表。当所需的表被移除时,94.6%的正确预测恰好命名了其中一个表,这与从参数化知识中重建缺失模式的现象一致。在冻结金标准条件下,对于两个主要求解器,三种表示方式之间的差异最多为2个百分点,最宽的配对95%置信区间将差异限制在±4个百分点以内。我们使用来自不同系列的一个推理模型观察到了相同的定性模式。当检索受覆盖范围限制时,执行准确率对模式预算的敏感度高于对测试的序列化方式的敏感度。该诊断基准以及用于构建和评估它的代码均可公开获取。
英文摘要
Data agents over structured sources must fit database schema into the model's context window. Large catalogs can span many databases and thousands of columns, so cost constraints may require choosing between table coverage and serialization detail well before the context window is full. We introduce BudgetSchemaBench, an execution-grounded diagnostic for this setting. Its construction derives relevance labels mechanically from gold SQL, without human- or LLM-authored ground truth. Using a pooled 80-database catalog, we sweep four schema-context budgets and compare three representations while keeping each retriever's table ranking fixed. A source-namespace check rejects queries that obtain the correct result from the wrong database. The evaluation covers three conditions: end-to-end retrieval; frozen-gold, in which the required tables are guaranteed; and a probe that removes those tables. For the primary solver with raw serialization, raising the budget from 2.5% to 50% of the catalog improves execution accuracy on 1,279 held-out questions by 18 percentage points under lexical retrieval but only 3 under dense retrieval; the dense retriever already finds most required tables at the smallest budget. When the required tables are removed, 94.6% of correct predictions name one of them exactly, consistent with reconstruction of absent schema from parametric knowledge. For the two main solvers in the frozen-gold condition, the three representations differ by at most 2 percentage points, and the widest paired 95% confidence interval bounds the difference within +/-4 points. We observe the same qualitative patterns with one reasoning model from a different family. When retrieval is coverage-limited, execution accuracy is more sensitive to the schema budget than to the tested serializations. The diagnostic and the code used to construct and evaluate it are publicly available.