发表机构
Robotics Institute, Carnegie Mellon University(卡内基梅隆大学机器人研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ParticleSplat提出一种自监督的以对象为中心的3D潜在粒子泼溅方法,通过将多视图编码为3D高斯粒子来克服DLP的2D限制,实现无监督对象掩码和可控3D编辑,并提升机器人操作性能。
AI 中文摘要
我们提出了ParticleSplat,一种自监督的以对象为中心的学习方法,通过前馈3D高斯泼溅将场景分解为一组表示语义实体的潜在“粒子”。基于Deep Latent Particles (DLP)框架(该框架将图像表示为具有位置、尺度和视觉外观等属性的一组粒子),我们解决了DLP的一个关键局限性:其固有的2D性质,这阻碍了显式的3D空间和几何推理,而这些对于机器人操作等下游任务至关重要。利用潜在粒子与3D高斯原语之间的结构相似性,我们引入了一个通过新颖视图合成目标训练的3D潜在粒子空间。我们的模型将带有相机位姿的多视图联合编码到一个共享的3D以对象为中心的潜在空间中,然后将粒子转换为与粒子对齐的3D高斯,其组合重建整个场景。在模拟和真实世界数据集上,我们表明这种公式无需监督即可学习对象掩码,并支持可控的3D场景编辑,例如通过在潜在空间中修改粒子来移动对象。我们进一步证实,学习到的3D表示提高了机器人操作任务的下游性能。
英文摘要
We present ParticleSplat, a self-supervised object-centric learning method that decomposes scenes into a set of latent ''particles'' representing semantic entities through feedforward 3D Gaussian Splatting. Building on the Deep Latent Particles (DLP) framework, which represents images as a set of particles with attributes such as position, scale, and visual appearance, we address a key limitation of DLP: its inherently 2D nature, which prevents explicit 3D spatial and geometric reasoning that are critical for downstream tasks such as robotic manipulation. Leveraging the structural similarity between latent particles and 3D Gaussian primitives, we introduce a 3D latent particle space trained with a novel view synthesis objective. Our model jointly encodes multiple views with camera poses into a shared 3D object-centric latent space, then transforms particles into particle-aligned 3D Gaussians whose composition reconstructs the full scene. On simulated and real-world datasets, we show that this formulation inherently learns object masks without supervision and supports controllable 3D scene editing, such as moving objects by modifying particles in the latent space. We further establish that the learned 3D representation improves downstream performance on robotic manipulation tasks.
CommentsProject page: https://lyuxinghe.github.io/ParticleSplat-website/