发表机构
New York University(纽约大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对数据集搜索中连接发现问题,提出MosaicJoin方法,通过新颖草图策略平衡值级比较方法和列级方法的权衡,实现高扩展性,实验证明其性能优越,无需训练或微调,能处理大数据量。
AI 中文摘要
连接发现是数据集搜索中的核心任务,早期方法聚焦等值连接,而如今数据湖等中存在语义可连接但句法表示不同的列。近期方法面临权衡:值级比较方法能准确识别可连接列但对高基数列扩展性差;列级方法高效但无法捕捉细粒度值对齐。本文提出MosaicJoin,通过新颖草图策略平衡权衡,实现扩展性。实验表明其在所有基准测试中优于先前方法,速度快达66倍,无需训练或微调,对含多达57K值的查询列和含多达1M值的数据湖列都能稳健扩展。
英文摘要
Join discovery is a core task in dataset search, enabling users to find columns that can be joined with a given query column. Early approaches focused on equi-joins, but data lakes and open-data repositories often contain columns whose values refer to the same entity but use different syntactic representations. To address this challenge, recent approaches discover semantically joinable columns but face a fundamental trade-off: methods that perform value-level comparisons accurately identify joinable columns but scale poorly to columns with high cardinality; column-level methods that encode an entire column into a single embedding are efficient but do not capture the fine-grained value alignment that determines whether a join is possible. We present MosaicJoin, a value-level semantic join discovery method that balances this trade-off. MosaicJoin achieves scalability through a novel sketching strategy that approximates the joinability of a column pair without having to compare all values. At query time, MosaicJoin scores each candidate sketch using a joinability score at a cost bounded by the sketch size, making retrieval efficient even for high-cardinality columns. A query subsampling operator further reduces online search time with provable accuracy guarantees, enabling robust retrieval for large query columns. Extensive experiments show that MosaicJoin outperforms previously published methods across all benchmarks while running up to 66 times faster than other value-level methods. MosaicJoin requires no training or fine-tuning, and it scales robustly to query columns containing up to 57K values and data lake columns containing up to 1M values.