发表机构
Technical University of Munich(慕尼黑工业大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出工具增强框架,将度量计算从VLM权重中移出,通过几何工具提升三维空间推理,在ReVSI-Bench三项任务上显著提升,并分离感知与推理误差。
AI 中文摘要
视觉语言模型(VLMs)能够很好地描述场景,但在度量三维结构(如绝对距离、物理尺寸或自我中心方向)的推理方面表现不佳。我们提出一个模块化、预测器无关的工具增强框架,为一个小型VLM(Qwen3.5-4B)配备几何工具:3D目标检测、度量深度估计以及用于距离、尺寸和方位的确定性求解器。每个目标在其自身最佳视角的相机坐标系中被检测,工具利用该坐标系的姿态将每个检测结果提升到一个共享的世界坐标系。将度量计算从模型权重中移出并放入显式求解器,在ReVSI-Bench的四个任务中的三个上带来了显著提升:使用强单目检测器(WildDet3D)时,绝对距离的平均相对精度(MRA)从0.46提升至0.74,相对距离从39.1%提升至67.4%,相对方向从低于随机水平的25.9%提升至73.4%。由于任何检测器都可以在工具接口后替换,将真实检测器与真实边界框进行比较可以将感知误差与推理误差分离:编排仅损失0.03 MRA。目标尺寸受检测器限制:工具在真实边界框上近乎精确(0.97),但最佳真实检测器仅勉强超过无工具基线(0.61对比0.58),因为尺寸直接读取自单目检测器出错的框范围。在没有预定义流程的情况下,模型已能自行正确排序工具,在四个任务中的三个上匹配脚本化流程。
英文摘要
Vision-Language Models (VLMs) describe scenes well but reason poorly about metric 3D structure such as absolute distances, physical sizes, or egocentric directions. We present a modular, predictor agnostic, tool-augmented framework that equips a small VLM (Qwen3.5-4B) with geometric tools: 3D object detection, metric depth estimation, and deterministic solvers for distance, size and bearing. Each object is detected in the camera frame of its own best view, and the tools use that frame's pose to lift every detection into one shared world frame. Moving metric computation out of the model's weights and into explicit solvers yields large gains on three of four ReVSI-Bench tasks: with a strong monocular detector (WildDet3D), absolute distance rises from 0.46 to 0.74 Mean Relative Accuracy (MRA), relative distance from 39.1% to 67.4%, and relative direction from a below-chance 25.9% to 73.4%. Because any detector can be swapped in behind the tool interface, comparing real detectors against ground-truth boxes separates perception error from reasoning error: orchestration costs only 0.03 MRA. Object size is bounded by the detector: the tools are near-exact on groundtruth boxes (0.97) yet the best real detector barely beats the no-tool baseline (0.61 vs. 0.58), because size reads straight off a box extent monocular detectors get wrong. Without a predefined recipe, the model already sequences the tools correctly on its own, matching a scripted pipeline on three of four tasks.