AI 中文总结
研究如何让自主数据代理在表格数据湖解析查询,提出GRAFT方法。将表格检索视为图匹配问题,引入IGMS奖励,把子图生成重述为马尔可夫决策过程并学习价值函数,设计在线管道修剪候选空间,在相关数据集上效果优于基线。
AI 中文摘要
自主数据代理通过在表格数据湖中检索和推理证据来解析分析查询。现有方法独立地根据查询对表格评分,忽略了将它们联系起来的可连接性和可并性,返回下游代理无法整合的碎片化证据。我们提出了GRAFT(表格的图匹配检索与融合),主要有两个贡献。首先,我们将表格检索视为查询派生意图图与异构数据湖图之间的图匹配问题,并引入IGMS,一种在单一目标中耦合语义相关性、结构兼容性和证据多样性的对数行列式奖励。其次,我们将子图生成重新表述为马尔可夫决策过程,并通过对由反同态的规范压缩算子生成的自生成轨迹进行隐式Q学习来学习价值函数。我们还设计了一个三阶段在线管道,利用锚点可达性、谓词可接受性和奖励单调性在精确的IGMS评估之前大幅修剪候选空间。在适用于表格数据湖设置的Spider和BIRD上,GRAFT在逐点、贪婪扩展和结构感知基线中实现了最佳的召回率、精确率、F1和充分性,F1相对于最强基线有7.8%的相对增益,充分性有10.6%的相对增益,同时保持了高搜索效率。
英文摘要
Autonomous data agents resolve analytical queries by retrieving and reasoning over evidence in tabular data lakes. Existing methods score tables independently against the query and ignore the joinability and unionability that link them, returning fragmented evidence that downstream agents cannot integrate. We propose GRAFT (Graph-matched Retrieval and Fusion of Tables), structured around two principal contributions. First, we cast table retrieval as a graph matching problem between a query-derived intent graph and a heterogeneous data lake graph, and introduce IGMS, a log-determinant reward that couples semantic relevance, structural compatibility, and evidence diversity in a single objective. Second, we recast subgraph generation as a Markov decision process and learn a value function via implicit Q-learning on self-generated trajectories produced by a canonical compression operator that inverts the homomorphism. We further design a three-stage online pipeline that exploits anchor reachability, predicate admissibility, and reward monotonicity to greatly prune the candidate space before exact IGMS evaluation. On Spider and BIRD adapted to the tabular data lake setting, GRAFT achieves the best Recall, Precision, F1, and Sufficiency among point-wise, greedy-expansion, and structure-aware baselines, with relative gains of 7.8% in F1 and 10.6% in Sufficiency over the strongest baseline, while maintaining high search efficiency.