发表机构
Arizona State University(亚利桑那州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
TabJoinBench是一个用于评估语义、关系和混合数据湖场景中可连接表发现方法的基准,通过可组合扰动构建查询-候选对并保留可靠真实标注,支持可复现比较。
AI 中文摘要
连接发现旨在从大型数据存储库中识别能够以补充信息增强查询表的表,从而支持数据探索、特征工程和商业智能等下游任务。尽管已提出众多连接发现方法,但现有研究依赖于特定方法的基准构建,使得可复现且公平的比较变得困难。我们提出TabJoinBench,一个用于在语义、关系和混合数据湖场景中评估连接发现方法的基准。TabJoinBench使用特定于源的验证策略构建查询-候选对,通过可组合的扰动系统地引入结构、表示和语义变化,同时保留可靠的真实标注。我们评估了涵盖基于集合、基于特征和学习方法的代表性连接发现方法,以及通用语言模型嵌入基线,并公开发布处理后的数据集、真实标注和生成流程,以促进可复现的评估和未来研究。
英文摘要
Join discovery aims to identify tables from large data repositories that can augment a query table with complementary information, enabling downstream tasks such as data exploration, feature engineering, and business intelligence. Although numerous join discovery methods have been proposed, existing studies rely on method-specific benchmark construction, making reproducible and fair comparison difficult. We present TabJoinBench, a benchmark for evaluating join discovery methods across semantic, relational, and hybrid data lake scenarios. TabJoinBench constructs query-candidate pairs using source-specific validation strategies, systematically introduces structural, representation, and semantic changes through composable perturbations while preserving reliable ground truth. We evaluate representative join discovery methods spanning set-based, feature-based, and learned approaches, together with general-purpose language-model embedding baselines, and publicly release the processed datasets, ground-truth annotations, and generation pipeline to facilitate reproducible evaluation and future research.
Comments13 pages, 8 Tables, 1 Figure