AI 中文总结
CityToolVQA 通过外部视觉-几何工具链处理定量问题,提升 VLM 在城市低空环境中的三维空间认知,在 Open3D-VQA-v2 上显著提高准确率。
AI 中文摘要
CityToolVQA 针对视觉语言模型(VLM)在城市低空视觉问答中定量任务表现较弱的问题。我们将七项任务分为定性组和定量组:定性问题由 VLM 直接回答,而定量问题则由外部视觉-几何工具链处理,该工具链执行目标定位、分割、深度反投影和空间计算。该工具链可以以零样本方式附加到不同的 VLM 上;当主链检测无效或不可靠时,会触发深度辅助提示推理(DAPI)回退机制,且 CityToolVQA-SFT 将 8B 骨干模型适配到工具条件输入。在包含 73,324 个问题的 Open3D-VQA-v2 测试集上,CityToolVQA-SFT(Qwen3-VL-8B)达到了 67.6% 的整体准确率,将工具链附加到十个开源 VLM 上可将定量任务准确率提高 12.7 至 36.9 个百分点。这些结果表明,将显式的三维几何计算外部化,有效补充了仅依赖 RGB 的 VLM 在估计度量距离和物体大小方面的有限能力。
英文摘要
CityToolVQA addresses the weak performance of Vision-Language Models (VLMs) on quantitative tasks in urban low-altitude visual question answering. We divide the seven tasks into qualitative and quantitative groups: qualitative questions are answered directly by the VLM, whereas quantitative questions are handled by an external visual-geometric toolchain that performs object grounding, segmentation, depth back-projection, and spatial computation. The toolchain can be attached to different VLMs in a zero-shot manner; a Depth-Assisted Prompt Inference (DAPI) fallback is triggered when the main-chain detection is invalid or unreliable, and CityToolVQA-SFT adapts the 8B backbone to tool-conditioned inputs. On the 73,324-question Open3D-VQA-v2 test set, CityToolVQA-SFT (Qwen3-VL-8B) reaches 67.6% overall accuracy, and attaching the toolchain to ten open-source VLMs improves quantitative-task accuracy by 12.7-36.9 percentage points. These results indicate that externalizing explicit 3D geometric computation effectively complements the limited ability of RGB-only VLMs to estimate metric distances and object sizes.
Comments10 pages, 3 figures, 3 tables