通过代数哈希连接和大规模离散全球网格加速多边形内点谓词
Accelerating Point-in-Polygon Predicates via Algebraic Hash-Joins and Discrete Global Grids at Scale: An interactive benchmark
浏览论文内容
中文总结 AI 辅助
研究传统矢量多边形内点查询对海量数据扩展性差的问题,核心方法是用DuckDB对四种DGGS实现进行实证评估,主要贡献是证明数据预索引时DGGS能实现亚秒级连接延迟,释放执行引擎吞吐量。
中文摘要 AI 辅助
传统基于向量的多边形内点查询依赖计算昂贵的几何谓词,对海量数据集扩展性差,即便有空间索引加速。离散全球网格系统(DGGS)通过将几何离散为分层单元,把复杂空间关系转化为常数时间关系哈希连接提供了可扩展替代方案。但采用DGGS会带来数据编码开销,当前网格实现存在显著性能“工具差距”。本演示展示了一个交互式仪表板,使用DuckDB对四种DGGS实现(H3、S2、A5和ISEA4H)的计算权衡进行实证评估。通过渐进场景,该平台可视化实时编码开销,并展示预索引空间数据集如何消除此开销。最终证明,当数据预索引时,所有DGGS无论其数学复杂性或工具如何,都能收敛到亚秒级连接延迟,释放现代矢量化执行引擎的吞吐量。
英文摘要
Traditional vector-based point-in-polygon queries rely on computationally expensive geometric predicates that scale poorly for massive datasets, even when accelerated by spatial indices. Discrete Global Grid Systems (DGGS) offer a scalable alternative by discretizing geometries into hierarchical cells, transforming complex spatial relations into constant-time relational hash-joins. However, adopting a DGGS introduces an overhead to encode data, and current grid implementations exhibit a significant performance ``tooling gap.'' In this demonstration, we present an interactive dashboard that empirically evaluates these computational tradeoffs across four DGGS implementations (H3, S2, A5, and ISEA4H) using DuckDB. Through progressive scenarios, the platform visualizes the overhead of on-the-fly encoding and demonstrates how pre-indexing spatial datasets eliminates this overhead. Ultimately, the demo proves that when data is pre-indexed, all DGGS regardless of their mathematical complexity or tooling converge to sub-second join latencies, unlocking the throughput of modern vectorized execution engines.