arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

视觉提示就足够了?研究渐进式视觉支架下视觉语言模型的空间推理能力

Is Visual Prompting All You Need? Studying VLM Spatial Reasoning under Progressive Visual Scaffolds

Lars Benedikt Kaesberg, Tianyu Yang, Florian Valentin Wunderlich, Terry Ruas, Daniel Kurzawe, Jan Philip Wahle, Bela Gipp

arXiv 2608.21170首次发表:更新:

AI 中文总结

该研究针对 VLMs 的空间推理问题,在 SPaRC 基准中引入轻量视觉支架,可提升 VLMs 任务准确率并补充 GRPO 训练,揭示视觉呈现对 VLM 基准测量属性的关键作用。

AI 中文摘要

视觉语言模型(VLMs)在多模态推理领域发展迅速,但近期研究表明,其失败往往反映了视觉 grounding 与下游推理之间的相互作用。当潜在推理问题不变时,任务的视觉呈现如何影响模型性能及失败模式,目前仍不明确。我们在基于网格的视觉空间规划基准 SPaRC 中引入轻量输入侧支架,这些支架保留视觉模态,同时让空间结构更易获取。在多个 VLMs 上的实验显示,与原始视觉设置相比,这些支架可将任务准确率提升最多 34.0 个百分点,且能进一步补充基于 GRPO 的训练,相比原始视觉输入下近乎为零的增益,可额外提升最多 4.6 个准确率百分点。对端到端任务解决和目标检测的分析表明,这些增益与 grounding 相关错误的减少密切相关,而规则推理仍相对具有挑战性。我们发现,视觉呈现是决定 VLM 基准是测量 grounded 感知、下游推理还是两者混合的核心因素。

英文摘要

Vision-language models (VLMs) have advanced rapidly in multimodal reasoning, yet recent work shows that their failures often reflect an interaction between visual grounding and downstream reasoning. What remains less clear is how the visual presentation of a task shapes model performance and failure modes when the underlying reasoning problem is unchanged. We study this question in SPaRC, a benchmark for grid-based visual spatial planning, by introducing lightweight input-side scaffolds that preserve the visual modality while making spatial structure more accessible. Across multiple VLMs, these scaffolds improve task accuracy over the original visual setting by up to 34.0 percentage points and further complement GRPO-based training, yielding up to 4.6 additional accuracy points compared with near-zero gains on the original visual input. Analyses on both end-to-end task solving and object detection show that these gains are closely tied to reductions in grounding-related errors, while rule reasoning remains comparatively challenging. We find that visual presentation is a central factor that determines whether VLM benchmarks measure grounded perception, downstream reasoning, or a mixture of both.

CommentsAccepted at EMNLP 2026 (Findings)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑