发表机构
INSAIT, Sofia University “St. Kliment Ohridski”; Stanford University; Ulm University(索非亚大学圣克莱门特奥赫里德大学INSAIT研究所; 斯坦福大学; 乌尔姆大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
GraphWrit3R提出端到端三维场景图生成方法,输入点云或高斯泼溅,直接输出结构化JSON场景图,避免多阶段流水线、真实标注依赖和专有模型,在3DSSG基准上达到最先进性能。
AI 中文摘要
三维场景图通过编码物体、物体的语义属性以及物体之间的空间和功能关系,为复杂环境提供了一种结构化表示。当前的三维场景图生成方法存在若干根本性局限。它们依赖具有显式中间表示的复杂多阶段流水线,使系统脆弱且容易产生错误传播。它们假设推理期间可访问真实物体标注,这偏离了现实世界场景。它们依赖专有模型,阻碍了开源部署,或导致推理速度慢得令人望而却步。我们提出了GraphWrit3R,一种简单的端到端方法,它以三维点云、高斯泼溅或两者的组合作为输入,并直接输出一个完整的场景图作为结构化JSON脚本。该图列出了所有物体、它们的语义属性以及它们之间的关系,同时避免了上述所有局限。选择多种输入模态纯粹是为了通用性,允许一组权重处理不同场景。点云输入通过Sonata编码,高斯泼溅输入通过Chorus编码,两种模态都被投影到共享体素网格上,并通过一种新颖的逐体素对比对齐损失进行融合,然后由大型语言模型解码。作为LLM的自然结果,GraphWrit3R还支持开放词汇查询。在3DSSG基准上,我们的方法在物体类别、谓词和三元组召回率上达到了最先进性能,优于依赖推理期间真实物体标注的方法。我们进一步提供了定性结果,并分析了不同的输入模态配置、对比损失公式和令牌融合策略。
英文摘要
3D scene graphs provide a structured representation of complex environments by encoding objects, their semantic attributes, and the spatial and functional relationships between them. Current approaches for 3D scene graph generation suffer from several fundamental limitations. They rely on complex multi-stage pipelines with explicit intermediate representations, making systems fragile and prone to error propagation. They assume access to ground-truth object annotations during inference, which deviates from real-world scenarios. They depend on proprietary models, hindering open-source deployment, or incur prohibitively slow inference. We present GraphWrit3R, a simple end-to-end method that takes a 3D point cloud, Gaussian Splats, or a combination of both as input, and directly outputs a complete scene graph as a structured JSON script. The graph lists all objects, their semantic attributes, and the relationships between them, while avoiding all of the above mentioned limitations. The choice of multiple input modalities is purely for versatility, allowing a single set of weights to handle diverse scenarios. Point cloud inputs are encoded via Sonata and Gaussian Splat inputs via Chorus, with both modalities projected onto a shared voxel grid and fused through a novel per-voxel contrastive alignment loss before being decoded by a large language model. As a natural consequence of the LLM, GraphWrit3R also supports open-vocabulary querying. On the 3DSSG benchmark, our method achieves state-of-the-art performance on object class, predicate, and triplet recall, outperforming methods that rely on ground-truth object annotations during inference. We further provide qualitative results and analyze different input modality configurations, contrastive loss formulations, and token fusion strategies.
CommentsProject page at https://graphwrit3r.insait.ai