面向神经纹理压缩的线程高效解码
Thread-Efficient Decoding for Neural Texture Compression
- Advanced Micro Devices, Inc.(超威半导体公司)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对神经纹理压缩存在的GPU线程发散问题,提出共享解码器MLP架构结合渐进式解码器冻结与CLIP语义聚类的方法,在保持渲染质量的同时降低线程发散25%-52%,在Radeon RX 9070 XT上实现最高8.48倍加速。
AI中文摘要:
神经纹理压缩(NTC)的压缩率高于BCn格式,但存在GPU线程发散问题,大幅降低运行时性能。本研究提出共享解码器MLP架构,采用渐进式解码器冻结策略训练,结合纹理聚类,使线程发散降低25%-52%,同时保持渲染质量。在500余种纹理及多个真实渲染场景上评估,与非共享基线相比,在Radeon RX 9070 XT GPU上实现最高8.48倍加速。核心贡献包括:(1)统一共享解码器架构,通过纹理分组减少发散;(2)渐进式解码器冻结的训练方案,提升稳定性与重建精度;(3)基于CLIP嵌入的语义聚类策略,分组相似纹理以实现高效解码器共享;(4)全面的性能与 ablation 研究验证方法有效性。
英文摘要:
Neural texture compression (NTC) achieves higher compression ratios than BCn formats but suffers from GPU thread divergence, which significantly reduces runtime performance. In this work, we propose a shared decoder MLP architecture -- trained with a gradual decoder freezing schedule -- combined with texture clustering to reduce thread divergence by 25%-52% while preserving rendering quality. We evaluate our method on over 500 textures and multiple real rendering scenes, demonstrating up to 8.48x speedup on the Radeon RX 9070 XT GPU compared to non-shared baselines. Our key contributions include: (1) a unified shared decoder architecture that reduces divergence by grouping textures; (2) a training recipe with gradual decoder freezing that improves stability and reconstruction accuracy; (3) a semantic clustering strategy using CLIP embeddings that groups similar textures for effective decoder sharing; and (4) comprehensive performance and ablation studies validating our approach.