arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

可直接操作的SIMD位切片:一种用于内存高效谓词评估的框架

Direct-Operable SIMD Bit-Slicing: A Framework for Memory-Efficient Predicate Evaluation

Arunkumar Mathiyazhagan

arXiv 2608.26368首次发表:更新:

AI 中文总结

该研究提出基于Project Panama Vector API的SIMD位切片框架,可直接处理压缩数据流,大幅降低内存开销并提升谓词评估与查询处理的速度。

AI 中文摘要

传统Java对象模型因对象头和内部填充会产生大量内存开销,常成为数据密集型分布式系统的性能瓶颈。本文提出一种新型框架,利用Project Panama Vector API直接对位切片压缩数据流执行谓词评估。通过将标准行导向数据转置为并行位平面,该框架无需预先解压即可使用SIMD(单指令多数据)指令评估复杂过滤器,支持整数、长整型(时间戳)、双精度浮点数(通过IEEE 754保序变换)及字符串(通过字典编码)。基准测试显示,该框架内存占用最多降低8倍,同时保持或超过未压缩标准Java集合的吞吐量;在5000万行的TPCDS建模数据上进行端到端评估,5种典型的过滤器密集型查询模式相比标量扫描实现2.4至10.8倍加速,对TPCDS列的扩展类型基准测试显示,时间戳、十进制数及字典编码字符串的加速比为1.5至43倍。

英文摘要

Traditional Java object models introduce significant memory overhead due to object headers and internal padding, often leading to performance bottlenecks in data-intensive distributed systems. This paper presents a novel framework that utilizes the Project Panama Vector API to perform predicate evaluation directly over bit-sliced, compressed data streams. By transposing standard row-oriented data into parallel bit-planes, we demonstrate a mechanism to evaluate complex filters using SIMD (Single Instruction, Multiple Data) instructions without requiring prior decompression. The framework supports integers, longs (timestamps), doubles (via IEEE 754 order-preserving transformation), and strings (via dictionary encoding). Our benchmarks indicate a reduction in memory footprint by up to 8x while maintaining or exceeding the throughput of uncompressed standard Java collections. End-to-end evaluation on TPCDS-modeled data at 50M rows demonstrates 2.4-10.8x speedup over scalar scans across five representative filter-heavy query patterns, with extended type benchmarks on TPCDS columns showing 1.5-43x speedups for timestamps, decimals, and dictionary-encoded strings.

Comments25 pages, 8 figures, 12 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑