arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SceneScaffold:面向统一三维场景理解的活动场景状态构建

SceneScaffold: Active Scene-State Construction for Unified 3D Scene Understanding

Xiangqi Li, Libo Huang, Jiarui Zhao, Weilun Feng, Chuanguang Yang, Zhulin An, Yongjun Xu

arXiv 2609.33518首次发表:更新:

发表机构

Institute of Computing Technology, Chinese Academy of Sciences; University of Chinese Academy of Sciences(中国科学院计算技术研究所; 中国科学院大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

SceneScaffold提出活动场景状态构建框架,将视觉瓶颈转为主动场景组织者,构建角色感知空间脚手架,提升统一三维场景理解中关系密集与空间模糊场景的推理稳定性。

AI 中文摘要

近期三维大型多模态模型(3D-LMMs)依赖视觉瓶颈将复杂的三维场景证据压缩为与大型语言模型(LLMs)兼容的有限数量的视觉标记。然而,当前的视觉瓶颈往往被动地将异构的三维证据压缩成同质的、以对象为中心的标记序列,导致场景的空间组织未被充分表示。这种表示不足迫使LLM从扁平化的标记序列中恢复空间关系,导致在关系密集和空间模糊的场景中推理不稳定。为解决此问题,我们提出SceneScaffold,一个用于统一三维场景理解的活动场景状态构建框架。SceneScaffold将视觉瓶颈从被动的特征压缩器重新定义为主动的场景组织者,在语言推理之前构建一个角色感知的空间脚手架。具体而言,SceneScaffold将超点级视觉证据组织成具有不同结构角色的场景状态组件:实体状态保留核心对象语义,场景框架状态通过边界和区域锚点维持空间参考,关系状态编码对象-环境交互线索,全局摘要提供紧凑上下文。通过这种角色感知的构建,SceneScaffold在语言推理之前为LLM提供空间组织的场景表示。在统一三维场景理解任务(包括三维视觉定位、问答和密集描述)上的实验证明了SceneScaffold的有效性,而诊断结果进一步展示了其在关系密集和空间模糊案例中的适用性。代码可在以下https URL获取。

英文摘要

Recent 3D large multimodal models (3D-LMMs) rely on a visual bottleneck to compress complex 3D scene evidence into a limited number of visual tokens compatible with large language models (LLMs). Current visual bottlenecks, however, often passively compress heterogeneous 3D evidence into a homogeneous object-centric token sequence, leaving the spatial organization of the scene under-represented. This under-representation forces the LLM to recover spatial relations from a flattened token sequence, leading to unstable reasoning in relation-intensive and spatially ambiguous scenes. To address this issue, we propose SceneScaffold, an active scene-state construction framework for unified 3D scene understanding. SceneScaffold reformulates the visual bottleneck from a passive feature compressor into an active scene organizer, constructing a role-aware spatial scaffold before language reasoning. Specifically, SceneScaffold organizes superpoint-level visual evidence into scene-state components with distinct structural roles: entity states preserve core object semantics, scene-frame states maintain spatial references via boundary and region anchors, relation states encode object-environment interaction cues, and a global summary provides compact context. Through this role-aware construction, SceneScaffold provides the LLM with a spatially organized scene representation before language reasoning. Experiments on unified 3D scene understanding tasks, including 3D visual grounding, question answering, and dense captioning, demonstrate the effectiveness of SceneScaffold, while diagnostic results further show its applicability to relation-intensive and spatially ambiguous cases. Code is available at https://github.com/lixiangqi707/SceneScaffold.

CommentsAccepted by NeurIPS 2026 (Spotlight)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑