发表机构
Department of Statistics, Brigham Young University(杨百翰大学统计系)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出层次贝叶斯模型,将模式链接视为图上的结构化子集选择,通过成对耦合和数据库级随机效应,在BIRD和Spider上改善相关子图恢复与预测校准。
AI 中文摘要
将大语言模型接地到关系数据库需要选择与查询相关的表和连接。现有的模式链接方法通常独立地对表或列进行评分,尽管目标对象通常是外键图的连通子图。我们提出了一种层次贝叶斯模型,将模式链接视为小型属性图上的结构化子集选择。来自查询-表和查询-列特征的单一证据与对外键边的成对自逻辑耦合相结合,使得当弱信号表连接其他相关表时,可以选择它们。数据库级别的随机效应模拟基线包含率和耦合强度的变化,并且模式图的小尺寸允许对条件似然和包含概率进行精确评估,我们在参数后验的拉普拉斯近似上对其求平均。在基准数据集BIRD和Spider上,正图耦合恢复了独立评分遗漏的表,但将边际包含概率向上移动。联合估计截距和耦合恢复了校准,数据库级别的部分池化适应了跨模式的结构效应。图耦合后验为精确的相关子图分配了更多概率,并产生了具有接近名义覆盖的信息性预测集,而相应的独立模型的预测集覆盖不足。拟合模型还报告了每个数据库中图权重的后验。稳定性和均值偏移结果解释了耦合何时有助于恢复以及为什么截距必须补偿它。因此,我们主要将该方法评估为表子集上的后验预测分布,而不是作为最大化精度的选择器。
英文摘要
An LLM answering a question about a relational database must know which tables hold the information. Showing too few makes the question unanswerable, while showing too many wastes context. Most schema-linking methods score each table alone, though the required tables are usually connected by foreign keys and can include connectors matching the question only weakly. We develop a hierarchical Bayesian model that combines question evidence with the foreign-key graph and returns empirically near-calibrated probabilities over complete sets of tables, not a single list, so later stages can tell when a selection is uncertain. Schema graphs are small enough to evaluate every subset exactly. With a fixed context budget the graph improves the chance of retaining every required table, and the benefit grows with schema size, turning positive at about seven tables. When the model also chooses how many tables to keep, a learned prior on subset size accounts for the whole improvement and the graph adds nothing. Fit the size prior first, and use the graph when the budget is fixed. The model's predictive sets cover the true set about as often as promised, where independent scoring is overconfident.