发表机构
School of Biomedical Engineering and Technology, Tianjin Medical University; Medical School, Tianjin University(天津医科大学 biomedical 工程与技术学院; 天津大学医学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对工业编解码器高吞吐低压缩比与上下文混合算法高压缩比低速度的矛盾,提出自适应异构压缩架构GPX,结合结构决策探针、GPU变换与AVX2优化,在Silesia基准上实现高压缩比与高吞吐,并严格支配Zstandard Levels 10-14。
AI 中文摘要
现代无损数据压缩面临严峻的二元对立:工业流式编解码器(如Zstandard)优先考虑吞吐量(10至1000 MiB/s),但牺牲了压缩密度;而上下文混合算法虽能达到更高的压缩比,却以无法实际使用的串行速度(0.1至1 MiB/s)运行。我们提出GPX,一种面向结构化数据流的自适应异构压缩架构。尽管针对高吞吐企业级和科学大块负载(>= 4 MiB)进行了优化,GPX也能透明地处理任意流长度,直至亚千字节输入(N >= 256 B)。GPX结合了亚15微秒的原生结构决策探针、GPU原生可逆域变换、AVX2多假设序列优化以及自适应熵边界。在规范的Silesia基准(202.12 MiB)上,采用预注册的7次重复中位数+四分位距协议(W = 6个工作线程),GPX轨道A生成符合RFC 8878标准的.zst流,压缩至56.27 MiB,吞吐量为235.41 MiB/s,并具备原生线速解压(在未修改的libzstd上为3074.27 MiB/s),在评估的RFC 8878兼容配置中确立了非支配地位。GPX轨道B在259.76 MiB/s(置信区间[250.7, 287.6])下达到54.49 MiB,确立了经验帕累托前沿点,并支配了Stock Levels 8-14。GPX轨道A+B在191.90 MiB/s(置信区间[177.8, 206.7])下达到54.46 MiB,在非重叠95% Bootstrap置信区间下严格支配官方Stock Zstandard Levels 10-14(比Level 14小182.18 KB,比Levels 10-14快1.15倍至18.70倍)。零调参的样本外泛化在Canterbury(-0.650%)、Calgary(-0.124%)以及一个102.45 MB的现代编译二进制文件(-0.246%)上得到确认,而enwik8(+0.084%)则记录了分割器过度分割的真实失败模式。所有输出均通过100% SHA-256位精确往返验证。
英文摘要
Modern lossless compression enforces an acute dichotomy: industrial streaming codecs (e.g., Zstandard) prioritize throughput (10 to 1000 MiB/s) at the expense of ratio, while context-mixing algorithms achieve superior density at serial speeds (0.1 to 1 MiB/s). We present GPX, an adaptive multi-tier compression architecture for structured data streams. GPX operates transparently across arbitrary stream lengths (N >= 1,024 B, identity fallback below 1 KiB), combining an autonomous structural decision probe (6.7--486.2 us across 49 benchmark instances, median 31.6 us, <0.05% overhead), cache-tiled AVX2 SIMD reversible domain transforms, multi-hypothesis sequence optimization with in-flight frame MDL gating, and adaptive entropy boundaries. On canonical Silesia (202.12 MiB) under a 7-repeat protocol (W = 6 pinned P-cores), GPX Track A produces standard RFC 8878-compliant .zst streams compressing to 56.20 MiB at 191.44 MiB/s with native line-speed decompression (5,182.79 MiB/s at W=6, 1,706.52 MiB/s at W=1 on unmodified libzstd, representing 124.1% single-core speed of Stock Level 9). GPX Track B achieves 54.49 MiB at 299.23 MiB/s (CI [295.3, 305.7]) with 4,919.14 MiB/s decode, strictly dominating Stock Levels 9--14 across size and speed (1.04x faster and 1.95 MB smaller than Level 9; 151.78 KB smaller and 8.24x faster than Level 14). GPX Track A+B achieves 54.37 MiB at 204.87 MiB/s (CI [199.6, 210.3]) with 4,599.00 MiB/s decode, strictly dominating Levels 10--14 under 95% Bootstrap CIs (276.07 KB smaller than Level 14, 1.19x to 5.64x faster). Generalization is confirmed across 6 suites (49 instances, 48 streams): Canterbury (-0.65%), Calgary (-1.66%), PE binaries (-14.10%), and Transformer BFloat16 tensors (-10.61%, saving 3.28 MB vs Stock L9). Downstream SIMT nibble unpacking sustains 103.62 GiB/s (1.05x PCIe speedup). All outputs are bit-exact verified via SHA-256.
Comments10 pages, 1 figure, 3 tables. Camera-ready revision with comprehensive benchmark comparisons against SplitZip, ZipServ, and Brevis