发表机构
Huazhong University of Science and Technology; Harbin Engineering University; Agency for Science, Technology and Research (A*STAR)(华中科技大学; 哈尔滨工程大学; 新加坡科技研究局)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
M3SunAgent提出统一智能体,利用LLM规划空间视觉程序,协调工具完成单目度量深度估计与3D视觉定位,构建M3SI基准,性能优于现有模型。
AI 中文摘要
单目度量深度估计和3D视觉定位代表了单目3D空间理解(M3Sun)的两个互补基石,从中可以获取M3Sun所需的基本3D空间信息。然而,这些互补任务通常由独立的框架执行,这给具身智能系统带来了空间信息访问不灵活且不对齐的挑战。在本文中,我们提出了一个统一的单目3D空间理解智能体(M3SunAgent),它利用大语言模型(LLM)作为任务规划器进行空间视觉编程,灵活地生成结构化程序并协调工具。对于实例级度量深度估计任务,M3SunAgent调用目标检测器工具来定位目标,使用深度估计工具在选定点估计深度,并将这些预测聚合为实例级深度估计。我们还构建了M3Sun实例(M3SI)数据集,这是一个包含2,910个样本的基准,用于评估。对于单目3D视觉定位任务,M3SunAgent使用视觉-语言模型(VLM)工具来定位目标并输出基本空间属性,然后结合反投影工具与维度提升工具来预测其3D边界框。实验结果表明了M3SunAgent的优越性能。具体而言,在实例级单目度量深度估计的评估中,M3SunAgent在所有比较模型中取得了最佳性能,52.61%的预测实例分布在深度误差0.25(δ<0.25)以下。在单目3D视觉定位的评估中,M3SunAgent展现出比视觉和VLM模型整体更具竞争力的性能,达到了41.73%的3D平均交并比(mIoU),并超过最先进的MonoVLM模型3.62%。
英文摘要
Monocular metric depth estimation and 3D visual grounding represent the two complementary cornerstones of monocular 3D spatial understanding (M3Sun), from which the fundamental 3D spatial information required by M3Sun can be acquired. However, these complementary tasks are generally conducted by separate frameworks, which pose challenges of inflexible and unaligned spatial information access for embodied intelligence systems. In this paper, we propose a unified agent for monocular 3D spatial understanding (M3SunAgent) that leverages a large language model (LLM) as a task planner for spatial visual programming, which flexibly generate structured programs and coordinate tools. For instance-level metric depth estimation task, M3SunAgent invokes an object detector tool to locate the target, estimates depth at selected points with a depth estimation tool, and aggregates these predictions into an instance-level depth estimate. We also construct the M3Sun Instance (M3SI) dataset, a benchmark with 2,910 samples for evaluation. For monocular 3D visual grounding task, M3SunAgent uses a vision-language model (VLM) tool to locate the target and output basic spatial attributes, then combines back-projection tool with a dimension-lifting tool to predict its 3D bounding box. Experimental results demonstrate the superior performance of M3SunAgent. Specifically, in evaluations of instance-level monocular metric depth estimation, M3SunAgent achieves the best performance among all compared models, 52.61% of predicted instances are distributed below depth error 0.25 ($δ< 0.25$). In evaluations of monocular 3D visual grounding, M3SunAgent demonstrates overall competitive performance than vision and VLM models, reaching a 3D mean intersection over union (mIoU) of 41.73% and exceeding the state-of-the-art MonoVLM model by 3.62%.
Comments13 pages. 7 figures, submitted to IEEE Transactions on Circuits and Systems for Video Technology (TCSVT)