AI 中文总结
该研究提出无需训练的模块化框架ViewMind3D,将3D-QA分解为四个组件,在ScanQA和SQA3D上实现竞争力性能,证明通用LLMs与VLMs的模块化编排可实现有效3D推理。
AI 中文摘要
大型语言模型(LLMs)和视觉语言模型(VLMs)的最新进展为3D问答(3D-QA)开辟了新可能,3D-QA是具身AI和机器人感知的关键能力。然而,大多数现有方法依赖于特定于3D的训练或带有昂贵标注的微调,限制了其可扩展性和实际应用。我们提出ViewMind3D,这是一个完全无需训练的模块化框架,用于基于场景的多视图观测进行3D空间推理,无需完整的3D重建。该框架将3D-QA任务分解为四个可解释的组件:(1)问题驱动的多视图选择;(2)基于语言条件对象线索的引导视觉 grounding;(3)通过鸟瞰图(BEV)视角指示器进行空间上下文编码;(4)通过基于角色的推理生成结构化答案。这种设计无需模型微调即可实现结构化、鲁棒且可解释的推理。在ScanQA和SQA3D上的实验结果表明,ViewMind3D与先前的无需训练和微调的3D-LLMs相比具有竞争力的性能。特别是,我们的方法提高了空间接地问题类型的性能,例如SQA3D中的“What”问题,同时保持了50.8%的整体准确率,并在ScanQA上达到了73.4的CIDEr值。这些结果表明,通过对通用LLMs和VLMs进行模块化编排,可以在真实环境中实现机器人感知所需的有效3D推理。
英文摘要
Recent advances in large language models (LLMs) and vision-language models (VLMs) have enabled new possibilities for 3D question answering (3D-QA), a key capability for embodied AI and robotic perception. However, most existing methods rely on 3D-specific training or fine-tuning with costly annotations, limiting their scalability and real-world applicability. We present \textbf{ViewMind3D}, a fully training-free and modular framework for 3D spatial reasoning over multi-view observations of a scene without requiring complete 3D reconstruction. The framework decomposes the 3D-QA task into four interpretable components: (1) question-driven multi-view selection, (2) guided visual grounding with language-conditioned object cues, (3) spatial context encoding via a bird's-eye-view (BEV) viewpoint indicator, and (4) structured answer generation through role-based reasoning. This design enables structured, robust, and interpretable reasoning without requiring model tuning. Experimental results on ScanQA and SQA3D show that ViewMind3D achieves competitive performance compared to prior training-free and fine-tuned 3D-LLMs. In particular, our method improves performance on spatially grounded question types, such as ``What'' questions in SQA3D, while maintaining strong overall accuracy (50.8\%) and achieving 73.4 CIDEr on ScanQA. These results demonstrate that effective 3D reasoning can be achieved through modular orchestration of general-purpose LLMs and VLMs for robotic perception in real-world environments.