发表机构
Arizona State University; Adobe Research India(亚利桑那州立大学; Adobe 印度研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出JOINGR,一种利用数据库连接图进行表检索的方法,通过遍历连接边聚合分数,在BIRD、Spider和BEAVER基准上表现优异,并展现出跨域可迁移性。
AI 中文摘要
检索正确的表是对现实数据库进行文本到SQL查询的先决条件。密集表检索器独立地对模式元素进行排序,但这忽略了一个关键的证据来源:某些所需的表在问题中未被提及,只能通过它们与已相关表的连接关系来识别。我们引入了JOINGR,一种连接感知的表检索方法,它将数据库连接图视为检索空间。列被表示为图节点,而表内和外键关系被表示为类型化边。给定一个问题,JOINGR选择语义上相似的锚定表,使用查询条件下的评分器遍历连接边,并将由此产生的边贡献聚合为表分数。该评分器是一个轻量级MLP,基于冻结的查询、节点和边嵌入,并使用成对边际损失在黄金表上进行训练。在BIRD和Spider数据集上,JOINGR与最强的检索基线相当。在BEAVER(一个具有多跳表需求的挑战性企业基准)上,JOINGR在召回率上显著优于密集检索和重排序基线。跨域实验表明,所学到的评分器可以在基准之间转移,表明该方法捕获了可复用的连接图遍历行为。
英文摘要
Retrieving the right tables is a prerequisite for Text-to-SQL over realistic databases. Dense table retrievers rank schema elements independently, but this ignores a key source of evidence: some required tables are not mentioned in the question and become identifiable only through their join relationships to already relevant tables. We introduce JOINGR, a join-aware table retrieval method that treats the database join graph as the retrieval space. Columns are represented as graph nodes, while intra-table and foreign-key relationships are represented as typed edges. Given a question, JOINGR selects semantically similar anchor tables, traverses join edges with a query-conditioned scorer, and aggregates the resulting edge deposits into table scores. The scorer is a lightweight MLP on top of frozen query, node, and edge embeddings, trained with a pairwise margin loss over gold tables. On BIRD and Spider datasets, JOINGR is competitive with the strongest retrieval baselines. On BEAVER, a challenging enterprise benchmark with multi-hop table requirements, JOINGR substantially improves recall over dense retrieval and re-ranking baselines. Cross-domain experiments show that the learned scorer transfers across benchmarks, indicating that the method captures reusable joingraph traversal behavior.
Comments12 pages, 6 figures, 5 pages