发表机构
Peking University; North China Electric Power University; Hunan University(北京大学; 华北电力大学; 湖南大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对3D高斯溅射推理局限,本文提出CausalSplat框架,结合视觉语言模型与3D场景图,构建两个推理基准,在相关任务上实现最优性能与强泛化性。
AI 中文摘要
尽管3D高斯溅射(3DGS)已推动开放词汇场景理解的发展,但现有方法仍局限于显式查询,难以解释具身交互所需的隐式意图、复杂空间约束与常识推理。为填补这一空白,本文提出3D高斯分割推理任务,并构建两个基准数据集Causal-LERF与Causal-ScanNet,系统评估常识、空间、 affordance(功能可供性)及反事实推理能力。评估显示,当前最优方法在这些推理挑战上表现较差。因此,本文提出CausalSplat框架,将视觉语言模型与3D场景图结合,以解耦显式结构感知与隐式逻辑推理。大量实验表明,CausalSplat在本文的推理基准上达到最优性能,同时在标准指称与开放词汇3D分割任务上展现出强泛化性。项目页面:this https URL
英文摘要
While 3D Gaussian Splatting (3DGS) has advanced open vocabulary scene understanding, existing methods remain confined to explicit queries. They struggle to interpret implicit intents, complex spatial constraints, and commonsense reasoning required for practical embodied interactions. To address this gap, we introduce the task of reasoning 3D Gaussian segmentation and construct two benchmarks, Causal-LERF and Causal-ScanNet. These benchmarks systematically evaluate commonsense, spatial, affordance, and counterfactual reasoning. Evaluations reveal that current state of the art methods perform poorly on these reasoning challenges. Therefore, we propose CausalSplat, a framework that integrates vision-language models with 3D scene graphs to disentangle explicit structural perception from implicit logical inference. Extensive experiments demonstrate that CausalSplat achieves state of the art performance on our reasoning benchmarks while showing strong generalizability on standard referring and open vocabulary 3D segmentation tasks. Project Page: https://jiayuding031020.github.io/CausalSplat
CommentsAccepted to ACM MM 2026