arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

在图形处理器上渲染3D高斯分布

Rendering 3D Gaussians on a Graph Processor

Nicholas Fry, Ignacio Alzugaray, Mark Pupilli, Paul H. J. Kelly, Andrew J. Davison

arXiv 2607.15951首次发表:更新:

发表机构

Imperial College London Department of Computing(伦敦帝国学院计算机系)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究在IPU上实现3D高斯渲染器,输入为3D高斯地图,通过特定模式路由高斯基元,遵循BSP模型计算。评估了仅用SRAM实现的瓶颈及对性能和质量的影响,还探讨了对GPU和相关架构的意义。

AI 中文摘要

我们展示了在智能处理单元(IPU)上首次实现的3D高斯渲染器,该IPU由1472个仅带有片上静态随机存取存储器(SRAM)的独立瓦片组成,这些限制近似于高效传感器 - 处理器架构的属性。输入场景是来自真实世界序列的3D高斯地图。每个瓦片“拥有”帧缓冲区的一个屏幕空间区域;高斯基元通过在东北 - 西 - 南(NEWS)网格上的曼哈顿距离跳跃路由到目标瓦片,然后以扩展树模式分布到重叠邻居。计算遵循IPU的批量同步并行(BSP)模型,瓦片间通信在编译时定义。我们展示了这种硬件通过在核心之间实现本地数据传输来利用空间和时间局部性。我们评估了这种仅使用SRAM实现中的瓶颈:瓦片间带宽、每个瓦片的SRAM容量以及来自非均匀高斯密度的工作负载不平衡。我们分析了这些限制如何影响性能和渲染质量。这一探索为传统图形处理器(GPU)和3D表示提出了更广泛的问题,表明直接的流多处理器(SM)间通信可能为减少GPU内核中的动态随机存取存储器(DRAM)访问提供方法。我们讨论了这些对未来传感器上和无DRAM架构的影响。项目页面:此https URL

英文摘要

We present the first implementation of a 3D Gaussian renderer on an Intelligence Processing Unit (IPU), comprising 1,472 independent tiles with only on-chip SRAM; constraints that approximate properties of efficient sensor-processor architectures. Our input scenes are 3D Gaussian maps from real-world sequences. Each tile 'owns' a screen-space region of the framebuffer; Gaussian primitives are routed to destination tiles via Manhattan-distance hops on a north-east-west-south (NEWS) grid, then distributed to overlapping neighbours in an expanding tree pattern. Computation follows the IPU's Bulk Synchronous Parallel (BSP) model, with inter-tile communication defined at compile time. We show this hardware allows us to exploit spatial and temporal locality by enabling local data transfer between cores. We evaluate the bottlenecks in this SRAM-only implementation: inter-tile bandwidth, per-tile SRAM capacity, and workload imbalance from non-uniform Gaussian density. We analyse how these constraints affect performance and render quality. This exploration raises broader questions for conventional GPUs and 3D representations, suggesting that direct inter-SM (streaming multiprocessor) communication might offer ways to reduce DRAM access in GPU kernels. We discuss these implications for the future of on-sensor and DRAM-free architectures. Project page: https://nmjfry.github.io/ipu-3dgs/

CommentsProject page: https://nmjfry.github.io/ipu-3dgs/

Journal refEurographics Symposium on Rendering (Symposium Track), 2026

DOI:10.2312/sr.20261017

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑