TileGS:用于高斯溅射光栅化的分块局部深度分箱
TileGS: Tile-Local Depth Binning for Gaussian Splatting Rasterization
- Aalto University(阿尔托大学)
- Nokia Technologies(诺基亚技术)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
TileGS针对3D高斯溅射光栅化的长分块范围问题,提出分块局部深度分箱方法,在保持渲染精度的同时,显著提升了光栅内核及端到端帧的处理速度。
AI中文摘要:
实时3D高斯溅射(3DGS)实现了高渲染质量,但标准光栅化仍会遍历全局排序的分块流,这会产生较长的每分块范围和繁重的几何属性流量。本文提出TileGS,一种高斯溅射的分块局部重组方法。TileGS将每个长分块范围转换为一系列较短的深度局部范围,按从前到后的顺序对这些范围进行光栅化,并在粗排序不足以匹配基线合成的地方应用选择性修复。在桌面和笔记本Ada GPU上的9个场景基准测试中,我们的默认No-GW(无几何写入)变体在RTX 4090上实现了平均1.44倍的光栅内核加速,在RTX 4090上平均端到端帧加速为1.069倍,在RTX 1000 Ada上平均端到端帧加速为1.094倍,超过了广泛使用的优化开源3DGS实现gsplat;同时,TileGS的输出与gsplat的输出在数值噪声范围内完全匹配(|ΔPSNR| < 0.001 dB,|ΔSSIM| < 0.001,|ΔLPIPS| < 0.001)。完整套件的RTX 4090 Nsight Compute分析显示,尽管TileGS的SM吞吐量更低、活动 warp 占用率更低且DRAM流量更高,但它仍更快,总SASS线程指令减少了1.26倍。源属性分析证实,几何属性主导了剩余的内存压力(占总光栅流量的85.8%,占额外扇区的88.6%)。综合这些指标,可得出结论:TileGS通过减少有效光栅遍历工作而非减少字节量、优化合并、提高占用率或直接减少测得的warp发散来提升光栅性能。
英文摘要:
Real-time 3D Gaussian Splatting (3DGS) achieves high rendering quality, but standard rasterization still traverses a globally sorted tile stream that creates long per-tile ranges and heavy geometry-attribute traffic. We present TileGS, a tile-local reorganization of Gaussian splatting. TileGS turns each long tile range into a sequence of shorter depth-local ranges, rasterizes those ranges in front-to-back order, and applies selective repair where coarse ordering is insufficient to match baseline compositing. Across a 9-scene benchmark on desktop and laptop Ada GPUs, our default No-GW (No Geometry-Write) variant delivers a mean 1.44x raster-kernel speedup on RTX 4090 and mean end-to-end frame speedups of 1.069x on RTX 4090 and 1.094x on RTX 1000 Ada over gsplat--a widely used optimized open-source 3DGS implementation--while matching the gsplat output up to numerical noise (|Delta PSNR| < 0.001 dB, |Delta SSIM| < 0.001, |Delta LPIPS| < 0.001). Full-suite RTX 4090 Nsight Compute profiling reveals TileGS is faster despite lower SM throughput, lower active-warp occupancy, and higher DRAM traffic, while total SASS thread instructions fall by 1.26x. Source-attributed profiling confirms that geometry attributes dominate the remaining memory pressure (85.8% of total raster traffic and 88.6% of excess sectors). Together, these counters support the interpretation that TileGS improves raster performance by reducing effective raster traversal work, rather than by reducing byte volume, improving coalescing, increasing occupancy, or directly reducing measured warp divergence.