AI 中文总结
FaCTz是首个基于GPU的带误差界有损压缩器,可保留科学向量场的所有临界点,吞吐量最高达60 GB/s,比cpSZ的多线程CPU实现快约640倍,还提供优化压缩比的推测逐点模式
AI 中文摘要
带误差界的有损压缩对于存储和传输大规模科学模拟产生的向量场数据至关重要。尽管该方法能通过用户指定的误差界限制数值失真,但无法保留场的拓扑结构:微小的可容许扰动可能会产生或消除临界点,而下游特征分析依赖于这些临界点。现有GPU压缩器吞吐量很高,但与拓扑无关;唯一具有可证明临界点保留能力的压缩器cpSZ在CPU上运行,其吞吐量远低于现代GPU系统的数据生成速率。我们发现,尽管保留临界点本质上是耦合的顺序约束,但可将其重新表述为独立的并行任务,可基于每个块或推测性地基于每个点实现。我们提出FaCTz,这是首个保证临界点保留的基于GPU的带误差界有损压缩器。FaCTz提供了针对吞吐量优化的块模式和针对压缩比优化的推测性逐点模式。在三个向量场数据集上,FaCTz保留了所有临界点,同时实现了高达60 GB/s的吞吐量,比cpSZ的多线程CPU实现快约两个数量级(最高达约640倍)。其推测模式相比面向吞吐量的模式,进一步将压缩比提高了约一倍。
英文摘要
Error-bounded lossy compression is essential for storing and transferring the vector-field data produced by large-scale scientific simulations. Although it enforces a user-specified error bound to limit numerical distortion, it does not preserve the field's topology: small admissible perturbations can create or eliminate critical points on which downstream feature analysis depends. Existing GPU compressors achieve high throughput but are topology-agnostic, whereas the only compressor with provable critical-point preservation (cpSZ) runs on the CPU at throughput far below the data-generation rates of modern GPU-based systems. We observe that, although preserving critical points is inherently a coupled and sequential constraint, it can be reformulated into independent parallel tasks, either on a per-block basis or, speculatively, on a per-point basis. We present FaCTz, the first GPU-based error-bounded lossy compressor that guarantees critical-point preservation. FaCTz provides a block-wise mode optimized for throughput and a speculative per-point mode optimized for compression ratio. Across three vector-field datasets, FaCTz preserves every critical point while achieving throughput of up to 60 GB/s, approximately two orders of magnitude (up to approximately 640x) faster than the multithreaded CPU implementation of cpSZ. Its speculative mode further improves the compression ratio by approximately a factor of two over the throughput-oriented mode.