发表机构
Oregon State University; Oak Ridge National Laboratory; University of Kentucky; University of Oregon; New Jersey Institute of Technology(俄勒冈州立大学; 橡树岭国家实验室; 肯塔基大学; 俄勒冈大学; 新泽西理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出自适应渐进式压缩框架,通过自适应插值、系数分解等技术,在相同误差容差或PSNR下大幅提升压缩比,传输512GB数据时加速最高1.26倍,还可实现更高可视化质量。
AI 中文摘要
百亿亿次模拟生成数据的速度远超存储或分析能力,因此高效数据缩减至关重要。误差控制的有损压缩可在用户指定的误差界下实现高压缩比,但目标容差必须在压缩时固定。渐进式压缩放宽了这一限制,不过现有方法仍依赖固定的重构策略,且未充分利用分解系数间的相关性,限制了渐进式检索的效率。本研究提出一种自适应渐进式压缩框架,可提升误差界和峰值信噪比这两类常见目标的检索效率。研究贡献包括四点:一、针对不同目标,提出利用两种互补插值方案实现自适应渐进式压缩,并对其进行优化以达到高效率;二、提出系数分解这一新方法,利用去相关数据间常被忽略的空间相关性,进一步提升效率;三、开发自适应渐进式压缩工作流,可自动选择最适配的重构流水线并进行定制化优化;四、在五个真实世界科学数据集上,将所提框架与三种最先进的渐进式压缩器进行评估。实验结果显示,在相同请求误差容差下,所提框架的压缩比相较现有最优方法提升最高达42.3%;在相同峰值信噪比(PSNR)下提升最高达92.5%。在向远程站点传输512GB科学数据时,该框架在端到端数据传输性能上实现最高1.26倍的加速;此外,该方法在从存储中检索最少数据的同时,实现了最高的可视化质量。
英文摘要
Exascale simulations generate data far faster than it can be stored or analyzed, making efficient data reduction essential. Error-controlled lossy compression offers high compression ratios under user-specified error bounds, but the target tolerance must be fixed at compression time. Progressive compression relaxes this restriction, yet existing methods still rely on fixed refactoring strategies and do not fully exploit correlations among decomposed coefficients, limiting the efficiency of progressive retrieval. In this work, we present an adaptive progressive compression framework that improves retrieval efficiency for two common targets, namely error-bound and peak Signal-to-Noise ratios. Our contributions are fourfold. (1) We propose to leverage two complementary interpolation schemes for adaptive progressive compression toward different targets, and we optimize them to achieve high efficiency. (2) We propose coefficient decomposition, a novel method that exploits the commonly overlooked spatial correlations among decorrelated data, which further improves the efficiency. (3) We develop the adaptive progressive compression workflow with automatic selection of the best-fit refactoring pipeline and tailored optimizations. (4) We evaluate the proposed framework on five real-world scientific datasets against three state-of-the-art progressive compressors. Experimental results demonstrate that the proposed framework improves the compression ratio by up to $42.3\%$ under the same requested error tolerance and up to $92.5\%$ at the same PSNR, compared with the best-performing existing methods. When transferring $512$ GB of scientific data to remote sites, the framework delivers up to $1.26\times$ speedup in the end-to-end data transfer performance. Furthermore, our method achieves the highest visualization quality while retrieving the least amount of data from storage.