发表机构
AIST; University of Tsukuba; University of Technology Nuremberg; University of Oxford(日本国立 Advanced Industrial Science and Technology(AIST); 筑波大学; 纽伦堡工业大学; 牛津大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有3D-LLMs难以开展多物体间详细比较的问题,提出含MO3D数据集、Multi-3DLLM模型及Mini-apps基准的框架,该模型在MO3D任务上超越所有基线且对单物体分类有正向迁移
AI 中文摘要
我们解决三维大语言模型(3D-LLMs)中的一个根本性缺口:现有模型聚焦于单一物体/场景描述,难以开展详细的物体间比较。我们提出了一个用于多物体间详细物体级推理的框架,包含三个组件:(1)MO3D(三维中的多物体),一个需要细粒度多物体比较的指令数据集;(2)Multi-3DLLM,采用最小化的补丁交互Transformer(PIT),在保留局部几何的同时建模物体间/物体内关系;(3)Mini-apps,两个面向应用的基准测试(形状匹配、变更描述),用于探究几何理解的实际应用。近期的3D-LLMs和二维视觉语言模型(2D-VLMs)在这些任务上表现不佳,既缺乏以比较为中心的设计,也缺乏几何感知。相比之下,在我们的混合数据上训练的Multi-3DLLM学会了几何推理,在MO3D上超越了所有基线,并且对单一物体分类产生了正向迁移。
英文摘要
We address a fundamental gap in 3D-LLMs: existing models focus on single-object/scene description, struggling with detailed, inter-object comparison. We propose a framework for detailed object-level reasoning across multiple objects with three components: (1) MO3D (Multi-Object in 3D), an instruction dataset requiring fine-grained multi-object comparison; (2) Multi-3DLLM, using a minimal Patch-Interaction Transformer (PIT) that models inter-/intra-object relationships while preserving local geometry; (3) Mini-apps, two application-driven benchmarks (Shape Mating, Change Captioning) that probe geometric understanding for practical use. Recent 3D-LLMs and 2D-VLMs perform poorly on these tasks, lacking both comparison-centric design and geometric awareness. In contrast, Multi-3DLLM trained on our mixture data learns geometric reasoning, surpasses all baselines on MO3D, and provides positive transfer to single-object classification.
CommentsAccepted to CVPR 2026