发表机构
University of Florida(佛罗里达大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出CAESAR-LDAR压缩器,通过正交变换与自回归先验结合,在多元科学数据压缩中实现优异率失真性能,为多元科学压缩提供了实用设计原则。
AI 中文摘要
科学模拟会生成具有异质统计特性和依赖关系的物理场集合,但学习型压缩器通常会独立编码这些场,或依赖共享编码器,而未显式建模潜在空间中仍存在的结构。我们提出CAESAR-LDAR,这是一种误差可控的多元学习型压缩器,它在共享的CAESAR-V骨干网络基础上增加了两种互补机制:一种可训练的正交变换,用于重组对齐的潜在通道间的依赖关系;一种因果自回归分层先验,用于捕获变换后残留的局部空间结构。该变换通过矩阵指数参数化保持正交性,实现完全可逆,无需额外的惩罚项。所有变体均统一应用公共残差校正阶段,以强制执行要求的重建容差。在燃烧、气候和湍流数据上的实验表明,这两种机制在不同场景下发挥作用:当非线性编码器后仍存在大量线性跨通道依赖时,潜在去相关性作用最大;而当残留结构主要为局部或空间型时,自回归建模仍有效。二者结合在所有评估数据集上提供了最强或接近最强的率失真性能。全局变换带来的计算开销很小,而自回归编码则引入了更大的吞吐量权衡。更广泛地说,这些结果为多元科学压缩提供了实用设计原则:当潜在空间中存在可测量的全局跨通道依赖时,利用它;在更广泛的数据场景中,使用局部概率上下文作为互补机制。
英文摘要
Scientific simulations generate collections of physical fields with heterogeneous statistics and dependencies, yet learned compressors often encode those fields independently or rely on a shared encoder without explicitly modeling the structure that remains in latent space. We present CAESAR-LDAR, an error-controlled multivariate learned compressor that augments a shared CAESAR-V backbone with two complementary mechanisms: a trainable orthogonal transform that reorganizes dependence across aligned latent channels, and a causal autoregressive hierarchical prior that captures local spatial structure left after transformation. Orthogonality is maintained through a matrix-exponential parameterization, making the transform exactly invertible without an additional penalty. A common residual-correction stage is applied uniformly to all variants to enforce the requested reconstruction tolerance. Experiments across combustion, climate, and turbulence data show that the two mechanisms are useful in different regimes. Latent decorrelation helps most when substantial linear cross-channel dependence survives the nonlinear encoder, whereas autoregressive modeling remains effective when the remaining structure is primarily local or spatial. Their combination provides the strongest or near-strongest rate-distortion performance across the evaluated datasets. The global transform adds little computational overhead, while autoregressive coding introduces a larger throughput tradeoff. More broadly, the results suggest a practical design principle for multivariate scientific compression: exploit global cross-channel dependence when it is measurably present in latent space, and use local probabilistic context as a complementary mechanism across a wider range of data regimes.
Comments11 pages, 4 figures, 4 tables