解码-工作定律:基于边界的可证明精确空间连接压缩几何
The Decode-Work Law: Margin-Governed, Provably-Exact Spatial Joins over Compressed Geometry
浏览论文内容
中文总结 AI 辅助
提出基于Douglas-Peucker LOD梯度的可证明精确多边形相交连接,通过双边Hausdorff边界测试,实现比朴素解压再精炼少解码5.9倍顶点,并揭示解码工作量由符号间隙边界决定,与对象大小无关。
中文摘要 AI 辅助
过滤-精炼空间连接一直避免接触精确几何以获取认证候选对,但该领域从未对通过过滤的候选对的解压成本进行建模。当几何以压缩、渐进可解码的多分辨率编解码器存储时,连接的真实成本是解码的字节数。我们研究了在Douglas-Peucker细节层次(LOD)阶梯上可证明精确的多边形相交连接,通过双边Hausdorff边界测试进行认证,并做出两项贡献。首先,一个可复现的机制和工具:在真实的美国人口普查TIGER水域多边形上,我们的渐进证书连接返回精确的连接结果,同时解码的顶点数比朴素解压再精炼少3.4-16.8倍(中位数5.9倍),比Brinkhoff等人(1994)的单近似多步基线少约4.9倍,在31个工作负载中零正确性违规(与全精度预言机的集合相等性)。其次,我们称之为解码-工作定律的特征:解码工作量由每对符号间隙边界决定——即它距离谓词翻转边界的接近程度,与对象大小无关,因为证书仅下降到其分辨率超过边界为止。该定律在受控几何上清晰(保留R2=0.87,与大小无关),在真实数据上具有方向性(R2约0.55)。我们明确说明了不成立的情况:近边界顶点预测器是错误的模型(我们预先注册了一个并拒绝了它),选择性机制预测器未实现,最坏情况是在对抗性交错边界上的平凡Omega(v)读取界限。我们贡献了机制、预算诚实的解码核算和开放工具;我们不声称提出了新索引。
英文摘要
Filter-and-refine spatial joins have always avoided touching exact geometry for certified candidate pairs, but the field never modeled the decompression cost of the pairs that survive the filter. When geometry is stored in a compressed, progressively-decodable multiresolution codec, the join's true cost is bytes decoded. We study provably-exact polygon intersection joins over a Douglas-Peucker level-of-detail (LOD) ladder, certified by a two-sided Hausdorff-margin test, and make two contributions. First, a reproducible mechanism and harness: on real U.S. Census TIGER water polygons, our progressive certificate join returns the exact join result while decoding 3.4-16.8x (median 5.9x) fewer vertices than naive decompress-then-refine, and about 4.9x fewer than the single-approximation multi-step baseline of Brinkhoff et al. (1994), with zero correctness violations (set-equality against a full-precision oracle) across 31 workloads. Second, a characterization we call the decode-work law: decode work is governed by each pair's signed-clearance margin -- how close it is to the predicate-flip boundary -- independent of object size, because the certificate descends the ladder only until its resolution beats the margin. The law is clean on controlled geometry (held-out R2=0.87, size-independent) and directional on real data (R2 ~= 0.55). We are explicit about what does not hold: a near-boundary-vertex predictor is the wrong model (we pre-registered one and rejected it), a selectivity regime forecaster did not materialize, and the worst case is the trivial Omega(v) read bound on adversarially interleaved boundaries. We contribute the mechanism, budget-honest decode accounting, and an open harness; we do not claim a new index.