arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.15785cs.CV

RoofGS:基于Roofline模型的3D高斯溅射端到端加速

RoofGS: Roofline-Guided End-to-End Acceleration of 3D Gaussian Splatting

  • Harbin Institute of Technology(哈尔滨工业大学)

机构由 AI 辅助整理,请以论文原文为准。

Yang Luo, Yan Gong, Yongsheng Gao, Jie Zhao

AI总结:

该研究针对3D高斯溅射高分辨率GPU加速瓶颈,提出RoofGS框架,采用分辨率自适应量化深度排序键与范围感知比特级快速指数近似等优化,在RTX 4090的4K场景下实现10.1倍加速且PSNR损失极小。

AI中文摘要:

3D高斯溅射(3DGS)可实现实时新视图合成,但在GPU上高分辨率场景下仍存在局限。通过分阶段Roofline分析,我们确定两类不同的硬件瓶颈:前端受全局内存流量主导,而光栅化阶段受指令吞吐量限制。基于该分析,我们开发了RoofGS渲染框架,采用针对瓶颈的优化而非通用内核加速。对于内存受限的前端,我们设计了分辨率自适应的量化深度排序键,将每个键压缩至32位;对于计算受限的光栅化器,我们引入了范围感知的比特级快速指数近似,该近似适配不透明度剔除后的有界指数范围,并推导了逐像素误差边界。这两项核心技术辅以额外优化(内核融合、紧凑属性存储、剔除、双像素评估),进一步减少内存流量并提升指令级并行度。实验表明,在RTX 4090上4K分辨率下,RoofGS较3DGS实现10.1倍的端到端加速,吞吐量从61 FPS提升至616 FPS,仅带来0.028 dB的PSNR损失。

英文摘要:

3D Gaussian Splatting (3DGS) enables real-time novel-view synthesis but remains limited on GPUs at high resolutions. Through a stage-wise Roofline characterization, we identify two distinct hardware bottlenecks: global memory traffic dominates the front end, whereas instruction throughput limits rasterization. Guided by this analysis, we develop RoofGS, a rendering framework that applies bottleneck-specific optimizations rather than generic kernel acceleration. For the memory-bound front end, we design a resolution-adaptive quantized depth sorting key that compresses each key to 32 bits. For the compute-bound rasterizer, we introduce a range-aware bit-level fast exponential approximation tailored to the bounded exponent range after opacity culling, with a derived per-pixel error bound. These two core techniques are complemented by additional optimizations (kernel fusion, compact attribute storage, culling, dual-pixel evaluation) that additionally reduce memory traffic and improve instruction-level parallelism. Experiments show that RoofGS achieves a 10.1$\times$ end-to-end speedup over 3DGS at 4K on an RTX 4090, increasing throughput from 61 to 616 FPS, with only a 0.028 dB PSNR loss.

↑