发表机构
Institute of Automation, Chinese Academy of Sciences; University of Chinese Academy of Sciences; Shanghai Jiao Tong University(中国科学院自动化研究所; 中国科学院大学; 上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
DeGS是一种解耦工作负载解析与重组的可扩展3DGS架构,通过优化数据流提升PE利用率,在多场景多分辨率下较现有3DGS加速器实现吞吐量、加速比与能效提升,高PE扩展时仍保持高利用率。
AI 中文摘要
3D高斯溅射(3DGS)已成为实时新视图合成的领先技术,但现有3DGS加速器存在架构可扩展性差的问题:增加处理单元(PE)数量会导致渲染期间的性能提升微乎其微。我们发现根本原因是紧密耦合的“检查-同时融合”数据流,该数据流加剧了由不规则高斯覆盖导致的空间冗余,以及并行执行下异步逐像素终止导致的时间冗余所造成的PE利用率低下。为解决此问题,我们提出DeGS,一种用于高效3DGS推理的可扩展架构。为系统消除渲染固有的冗余,DeGS采用解耦数据流,将标准渲染过程中耦合的α检查、透射率检查和α融合,重组为连续的工作负载解析、重组和融合阶段。这使得碎片化、长度可变且依赖时间的工作负载可在融合前被重组为紧凑、无冲突且密集的工作负载,从而显著提高并行融合期间的PE利用率。采用28纳米工艺实现的DeGS,在不同场景和分辨率(720p至8K)下,相较于最先进的3DGS加速器(GSCore、GBU、GCC),实现了2.36倍至7.25倍的吞吐量、1.82倍至6.02倍的端到端加速比,以及1.59倍至4.42倍的能效提升。此外,从16个PE扩展至1024个PE时,DeGS在高分辨率下保持超过80%的PE利用率,显著优于现有加速器。
英文摘要
3D Gaussian Splatting (3DGS) has emerged as a leading technique for real-time novel view synthesis, yet existing 3DGS accelerators suffer from poor architectural scalability: increasing the number of PEs leads to marginal performance improvement during rendering. We identify that the root cause is the tightly coupled ``checking-while-blending'' dataflow, which exacerbates PE underutilization caused by spatial redundancy from irregular Gaussian coverage and temporal redundancy from asynchronous pixel-wise termination under parallel execution. To address this issue, we propose DeGS, a scalable architecture for efficient 3DGS inference. To systematically eliminate the redundancies inherent in rendering, DeGS exploits a decoupled dataflow, restructuring the coupled $α$-checking, transmittance checking, and $α$-blending of the standard rendering process into consecutive workload parsing, reorganization, and blending stages. This allows the fragmented, length-variable, and temporal-dependent workloads to be reorganized into compact, conflict-free, and dense workloads prior to blending, thereby significantly improving PE utilization during parallel blending. Implemented in 28 nm technology, DeGS achieves 2.36$\times$--7.25$\times$ throughput, 1.82$\times$--6.02$\times$ end-to-end speedup, and 1.59$\times$--4.42$\times$ energy efficiency over state-of-the-art 3DGS accelerators (GSCore, GBU, GCC) across diverse scenes and resolutions (720p to 8K). Moreover, scaling from 16 to 1024 PEs, DeGS maintains over 80\% PE utilization at high resolutions, significantly outperforming existing accelerators.
CommentsAccepted to the 59th IEEE/ACM International Symposium on Microarchitecture (MICRO 2026)