LoCoSplat:具有最小3D推理的实时前馈3D高斯泼溅
LoCoSplat: Real-Time Feed-Forward 3D Gaussian Splatting with Minimal 3D Reasoning
浏览论文内容
中文总结 AI 辅助
LoCoSplat利用局部平均代替重型3D网络,以0.14M参数MLP实现实时前馈3D高斯泼溅,在多个视图下超越先前方法,速度快4.2倍,内存减少6.7倍。
中文摘要 AI 辅助
前馈3D高斯泼溅(3DGS)越来越多地通过重型学习型3D网络聚合多视图证据。我们提出LoCoSplat(局部上下文泼溅),其动机源于观察到高斯是一种局部基元:一旦深度被预测,3D阶段必须添加的内容(尺度、旋转、不透明度)取决于每个锚点周围的点云,而该邻域的固定局部平均值足以提供这些信息,无需重型网络。LoCoSplat精确实现了这种平均值:它将点特征的16维线性投影泼溅到细网格和粗网格中,并在每个锚点处通过一个0.14M参数的逐点MLP读回两者;由于没有学习型3D网络和动态稀疏计算,其整个编码器作为一个fp16 CUDA图运行。在RealEstate10K上,LoCoSplat在6、12和24视图下,在PSNR、SSIM和LPIPS指标上均优于所有先前的前馈方法,且随着视图密度增加,优势进一步扩大(在24视图下,PSNR比先前体素对齐的最先进方法VolSplat高出3.3),并在对ACID的零样本迁移和对ScanNet的微调中进一步增长。它在一张NVIDIA RTX PRO 6000 GPU上以33毫秒重建6视图场景,是七种前馈方法中最快的,比先前最先进方法快4.2倍,训练速度快2.7倍(在24视图下快5.9倍),并且推理内存使用量减少6.7倍。
英文摘要
Feed-forward 3D Gaussian Splatting (3DGS) increasingly aggregates multi-view evidence with heavy learned 3D networks. We propose LoCoSplat (Local-Context Splatting), motivated by the observation that a Gaussian is a local primitive: once depth is predicted, what the 3D stage must add (scale, rotation, opacity) depends on the point cloud around each anchor, and a fixed local average of that neighbourhood is enough to supply it, no heavy network required. LoCoSplat realises exactly this average: it splats a 16-d linear projection of the point features into a fine and a coarse grid and reads both back at each anchor with a 0.14M-parameter pointwise MLP; with no learned 3D network and no dynamic sparse computation, its whole encoder runs as one fp16 CUDA graph. On RealEstate10K, LoCoSplat outperforms every prior feed-forward method on PSNR, SSIM, and LPIPS at 6, 12, and 24 views, with a margin that widens as views densify (+3.3 PSNR over VolSplat, the prior voxel-aligned state of the art, at 24 views) and grows further under zero-shot transfer to ACID and fine-tuning on ScanNet. It reconstructs a 6-view scene in 33 ms on one NVIDIA RTX PRO 6000 GPU, the fastest of seven feed-forward methods and $4.2\times$ faster than the previous state of the art, trains $2.7\times$ faster ($5.9\times$ at 24 views), and uses $6.7\times$ less inference memory.
发表机构
- Georgia Institute of Technology(佐治亚理工学院)
机构由 AI 辅助整理,请以论文原文为准。