arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DeGS:一种通过解耦工作负载解析与重组实现的可扩展3DGS架构

DeGS: A Scalable 3DGS Architecture via Decoupled Workload Parsing and Reorganization

Minnan Pei, Gang Li, Zeyu Zhu, Siting Wang, Junwen Si, Zhuoran Song, Yu Feng, Fangxin Liu, Xiaoyao Liang, Jian Cheng

arXiv 2608.02099首次发表:更新:

发表机构

Institute of Automation, Chinese Academy of Sciences; University of Chinese Academy of Sciences; Shanghai Jiao Tong University(中国科学院自动化研究所; 中国科学院大学; 上海交通大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

DeGS是一种解耦工作负载解析与重组的可扩展3DGS架构,通过优化数据流提升PE利用率,在多场景多分辨率下较现有3DGS加速器实现吞吐量、加速比与能效提升,高PE扩展时仍保持高利用率。

AI 中文摘要

3D高斯溅射(3DGS)已成为实时新视图合成的领先技术,但现有3DGS加速器存在架构可扩展性差的问题:增加处理单元(PE)数量会导致渲染期间的性能提升微乎其微。我们发现根本原因是紧密耦合的“检查-同时融合”数据流,该数据流加剧了由不规则高斯覆盖导致的空间冗余,以及并行执行下异步逐像素终止导致的时间冗余所造成的PE利用率低下。为解决此问题,我们提出DeGS,一种用于高效3DGS推理的可扩展架构。为系统消除渲染固有的冗余,DeGS采用解耦数据流,将标准渲染过程中耦合的α检查、透射率检查和α融合,重组为连续的工作负载解析、重组和融合阶段。这使得碎片化、长度可变且依赖时间的工作负载可在融合前被重组为紧凑、无冲突且密集的工作负载,从而显著提高并行融合期间的PE利用率。采用28纳米工艺实现的DeGS,在不同场景和分辨率(720p至8K)下,相较于最先进的3DGS加速器(GSCore、GBU、GCC),实现了2.36倍至7.25倍的吞吐量、1.82倍至6.02倍的端到端加速比,以及1.59倍至4.42倍的能效提升。此外,从16个PE扩展至1024个PE时,DeGS在高分辨率下保持超过80%的PE利用率,显著优于现有加速器。

英文摘要

3D Gaussian Splatting (3DGS) has emerged as a leading technique for real-time novel view synthesis, yet existing 3DGS accelerators suffer from poor architectural scalability: increasing the number of PEs leads to marginal performance improvement during rendering. We identify that the root cause is the tightly coupled ``checking-while-blending'' dataflow, which exacerbates PE underutilization caused by spatial redundancy from irregular Gaussian coverage and temporal redundancy from asynchronous pixel-wise termination under parallel execution. To address this issue, we propose DeGS, a scalable architecture for efficient 3DGS inference. To systematically eliminate the redundancies inherent in rendering, DeGS exploits a decoupled dataflow, restructuring the coupled $α$-checking, transmittance checking, and $α$-blending of the standard rendering process into consecutive workload parsing, reorganization, and blending stages. This allows the fragmented, length-variable, and temporal-dependent workloads to be reorganized into compact, conflict-free, and dense workloads prior to blending, thereby significantly improving PE utilization during parallel blending. Implemented in 28 nm technology, DeGS achieves 2.36$\times$--7.25$\times$ throughput, 1.82$\times$--6.02$\times$ end-to-end speedup, and 1.59$\times$--4.42$\times$ energy efficiency over state-of-the-art 3DGS accelerators (GSCore, GBU, GCC) across diverse scenes and resolutions (720p to 8K). Moreover, scaling from 16 to 1024 PEs, DeGS maintains over 80\% PE utilization at high resolutions, significantly outperforming existing accelerators.

CommentsAccepted to the 59th IEEE/ACM International Symposium on Microarchitecture (MICRO 2026)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑