arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Cadence:基于时间序列基础模型的带误差界的需求时间序列有损压缩

Cadence: Error-Bounded Lossy Compression of Demand Time Series with a Time-Series Foundation Model

Roberto Tacconelli

arXiv 2609.06008首次发表:更新:

AI 中文总结

Cadence利用330M参数时间序列基础模型与自适应算术编码器,实现带误差界的有损压缩,在多个真实数据集上显著优于经典方法,并揭示了基础模型在无损编码中的局限性。

AI 中文摘要

我们提出了Cadence,一种用于数值时间序列的带误差界的有损压缩器,它将一个330M参数的时间序列基础模型(Google TimesFM-3)与自适应算术编码器配对,保证每个样本满足$|\hat{x}_t-x_t|\le\tau$。一个负面结果限制了设计空间:对于无损编码,基础模型毫无价值,因为节省的比特数随预测器精度的对数变化,$\Delta b=\log_2(\mathrm{MAE_{old}}/\mathrm{MAE_{new}})$。因此,TimesFM-3相对于32抽头线性预测器的$1.51\times$优势,在20.28比特中仅节省0.60比特,中位数增益为+0.03%。带误差界的编码在一点上摆脱了这一限制:一旦预测落入误差带内,残差索引为零,样本几乎免费编码。Cadence的贡献包括:(1)一种具有上下文建模二值化的自适应区间编码器,在真实索引上比xz/zstd提升9.7%(15/15),并推翻了一个通用后端的发现;(2)一个确定性结果——预测在不同批量大小下并非比特级一致,且没有PyTorch配置能修复此问题,迫使将组大小和执行设备纳入容器格式;(3)对任何合理训练截止日期之后的数据集的领域局部化。在49个EIA-930平衡机构需求序列(2026年)上,Cadence比六种经典预测器中最佳者提升13.3%,在50个MTA客流序列(2026年)上提升28.3%:在297个序列-容差组合上中位数提升21.4%,赢得全部297个组合。与时间序列数据库用于保留的下采样相比,其保证的最坏情况误差在相同大小下严格$28$–$56\times$更紧。端到端地,一旦支付上下文引导成本,增益从六个月小时数据的6.8%到渐近的15.1%。为了在SDRBench上反驳领域声明,理论预测失败并应验:中位数-0.8%,27个组合中0个有增益。另外三个负面结果和八个撤回声明已完整报告。

英文摘要

We present Cadence, an error-bounded lossy compressor for numeric time series pairing a 330M-parameter time-series foundation model (Google TimesFM-3) with an adaptive arithmetic coder, guaranteeing $|\hat{x}_t-x_t|\leτ$ on every sample. One negative result constrains the design space: for lossless coding a foundation model is worth nothing, because bits saved are logarithmic in predictor accuracy, $Δb=\log_2(\mathrm{MAE_{old}}/\mathrm{MAE_{new}})$. So the $1.51\times$ advantage TimesFM-3 holds over a 32-tap linear predictor buys 0.60 bits of 20.28, a median gain of +0.03%. Error-bounded coding escapes this at one point: once a forecast lands inside the band the residual index is zero and the sample nearly free. Cadence contributes: (1) an adaptive range coder with context-modelled binarization, beating xz/zstd on real indices by 9.7% (15/15) and reversing a finding from a general-purpose back end; (2) a determinism result -- predictions are not bit-identical across batch sizes, and no PyTorch configuration repairs this, forcing group size and execution device into the container format; and (3) domain localization on corpora postdating any plausible training cutoff. On 49 EIA-930 balancing-authority demand series (2026) Cadence gains 13.3% over the best of six classical predictors, and 28.3% on 50 MTA ridership series (2026): 21.4% median over 297 series-tolerance pairs, winning all 297. Against downsampling, what time-series databases deploy for retention, its guaranteed worst-case error is $28$--$56\times$ tighter at equal size. End-to-end, once the context bootstrap is paid for, gains run from 6.8% at six months of hourly data to 15.1% asymptotically. Attempting to falsify the domain claim on SDRBench, theory predicts failure and delivers: -0.8% median, 0 of 27 pairs gaining. Three further negative results and eight retracted claims are reported in full.

Comments10 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑