arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

关于有损压缩的两个比特:宇宙学中压缩的极限

Two bits about lossy compression: On the limits of compression in cosmology

Hurum Maksora Tohfa, Matthew McQuinn

arXiv 2609.40239首次发表:更新:

发表机构

University of Washington(华盛顿大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究利用神经压缩器在宇宙学模拟和观测数据上实现比高斯最优编码器低9%-27%的比特率,且模型可跨数据集迁移,仅需1-6 bpp达到百分之一精度。

AI 中文摘要

天文学正处于大规模巡天和用于解释这些巡天的大型模拟套件的时代,这两者都存储成本高昂且共享缓慢。模拟输出以32位浮点数存储,然而数值噪声和天体物理不确定性使得优于百分之一水平的像素精度变得不必要。由于宇宙学场在统计上是均匀的,且模式振幅近似高斯分布,经典的率失真理论结果可直接适用,而非高斯结构允许进一步压缩。我们使用了科学压缩软件包SZ3,以及一种神经压缩器,该压缩器固定量化,并通过自回归Transformer学习每个量化区间的概率。SZ3仅在像素尺度上场平滑或数值在单点堆积时才优于高斯方法,而神经方法在我们考虑的每个场上都达到或优于SZ3,在固定失真下比高斯最优编码器低9%至27%,在非高斯性最强的场上优势最大。由于设定误差的是量化器而非网络,训练不佳的模型可能浪费比特但绝不会牺牲精度。仅在弱引力透镜收敛图上训练的模型无需重新训练即可迁移到N体密度场,这表明它学习了宇宙学结构的通用属性而非单个数据集的特征。欧拉网格在百分之一精度下仅需每像素1-4比特(bpp),而粒子位移和速度在其所需精度下需要5-6比特。该方法也适用于观测数据:在鲁宾天文台数据预览1的叠加图上,当量化步长设为背景噪声的四分之一时,微调后的Transformer需要4.0 bpp,比高斯编码器少27%。

英文摘要

Astronomy is in an era of enormous sky surveys and of the large simulation suites needed to interpret them, both costly to store and slow to share. Simulation outputs are stored as 32-bit floats, yet numerical noise and astrophysical uncertainties make better than percent-level pixel accuracy unnecessary. Because cosmological fields are statistically homogeneous with nearly Gaussian mode amplitudes, classic rate-distortion results apply directly, and non-Gaussian structure permits further compression. We use the scientific compression package SZ3, and a neural compressor that fixes the quantization and learns the probability of each quantization bin with an autoregressive transformer. SZ3 beats the Gaussian approach only where the field is smooth on the pixel scale or its values pile up at a single point, while the neural approach matches or beats SZ3 on every field we consider and is 9 to 27\% below the Gaussian-optimal coder at fixed distortion, with the largest margin where the field is most non-Gaussian. Because the quantizer, not the network, sets the error, a poorly trained model can waste bits but never cost accuracy. A model trained only on weak lensing convergence maps transfers to $N$-body density fields without retraining, suggesting it has learned generic properties of cosmological structure rather than features of one dataset. Eulerian grids need only 1-4 bits per pixel (bpp) at percent-level accuracy, and particle displacements and velocities 5-6 at the precisions they require. The approach carries over to observational data: on Rubin Observatory Data Preview 1 coadds, with the quantization step set to a quarter of the background noise, the fine-tuned transformer needs 4.0 bpp, 27\% fewer than the Gaussian coder.

Comments10 Pages, 1 figure

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑