arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VIREL:用于精确和有误差界的浮点时间序列压缩的格上验证整数残差编码

VIREL: Route-Local Lattice Residual Compression for Exact and Error-Bounded Floating-Point Time Series

Yue Zhang, Jiatao Lin, Haopeng Chen

arXiv 2607.22433首次发表:更新:

AI 中文总结

研究针对流浮点时间序列压缩问题,提出VIREL框架。通过纳入格、预测残差、处理混合精度等操作,实现精确和有误差界压缩。实验表明其在不同场景下压缩率高,整数域残差预测是主要增益源,还能跨核心扩展及作页面编解码器。

AI 中文摘要

流浮点数据通常在十进制或物理分辨率格上生成,但大多数精确压缩器仍直接对IEEE二进制表示进行建模。本文提出了VIREL,一个用于精确和有误差界的浮点时间序列压缩的可按页流处理的框架。VIREL首先将每个值纳入一个经过验证的十进制或误差格,然后在所得整数生成域而非IEEE字异或空间中预测残差。为处理混合精度和异常,VIREL将值路由到兼容通道并使预测器状态在每个通道局部化。其压缩优先的精确配置文件进一步应用了带成本的格步长归一化:对于仿射子格q = d z + r,仅当完整帧成本降低时才对紧凑坐标z进行编码。有误差界的配置文件以精确可除性规则应用相同思想,仅在重建相同格点且因此不消耗误差预算时存储q/d。在具有独立的1024值页面的规范精确流中,VIREL-Exact-Fast达到6.0243倍压缩,而VIREL-Exact-Upper达到7.0287倍,且比最强的评估先前精确编解码器发出的字节数少得多。在48个独立分页流中的7470万个值上,两个精确配置文件分别达到8.0629倍和9.6490倍。在15个Serf流上误差界为1e-3时,VIREL-EB达到12.1094倍,比最强的兼容外部基线少发射12.54%的比特,并保留所有检查的逐点界。消融实验表明整数域残差预测是增益的主要来源,而格步长归一化和路由持久多通道预测在混合分辨率下保持该增益。VIREL还可跨CPU核心扩展并作为Apache TsFile页面编解码器运行。

英文摘要

Floating-point page codecs exploit temporal smoothness, but existing methods keep prediction state in different representation domains: IEEE 754 words, erased IEEE 754 words, decimal fields, or integer surrogates. Which domain should carry temporal prediction state inside a database page remains an open question. We present VIREL, a page codec that predicts route-local lattice-coordinate residuals for values admitted to exact or error-bounded integer coordinates. Separate routes preserve the history of mixed source resolutions, and cost-based lattice-step normalization stores compact coordinates such as $z$ for $q=dz+r$ or $q/d$ for divisible error-lattice indices while restoring the same lattice point before reconstruction. On canonical exact streams with independent 1,024-value pages, VIREL-Exact-Fast reaches 6.0243$\times$ and VIREL-Exact-Upper reaches 7.0287$\times$, emitting 22.4% fewer bytes than the strongest evaluated exact baseline. On 74.70 million values in 48 streams, the two profiles reach 8.0629$\times$ and 9.6490$\times$. At $ε=10^{-3}$ on 15 Serf streams, VIREL-EB reaches 12.1094$\times$, emits 12.54% fewer bits than the strongest compliant error-bounded baseline, and preserves all pointwise bounds. Ablations show that integer-domain residual prediction and $q/d$ factoring reduce output by 53.05% and 18.75% in their respective settings. The Fast profile scales to 1,034/1,196 MB/s encode/decode at 64 cores. As an Apache TsFile codec, it writes 27.9-30.1% fewer complete-file bytes than DeXOR and ELF*, and with LZ4 reaches 105.80/102.32/110.30 MB/s on full-scan, range-scan, and aggregate queries, faster than the encoded baselines in all three read paths.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑