arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向决策关键型应用的多模态大语言模型中可解释且资源高效的空间推理

Explainable and Resource-Efficient Spatial Reasoning in Multimodal LLMs for Decision-Critical Applications

Piyush Jain, Kousik Dasgupta, Rajarshi Roy, Subarna Tripathi

arXiv 2607.27145首次发表:更新:

发表机构

Heritage Institute of Technology; Kalyani Government Engineering College; IAIRO; Intel Corporation(遗产技术学院; 卡利亚尼政府工程学院; IAIRO; 英特尔公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对多模态大语言模型空间判断不透明及细粒度空间理解不足的问题,提出无需训练的ByDeWay-V2框架,结合YOLO-World-L注入空间谓词,在多个基准上显著提升空间推理性能,适配资源受限场景。

AI 中文摘要

随着多模态大语言模型(MLLMs)越来越多地部署在机器人、具身AI和安全监控等决策关键型流程中,其空间判断的不透明性限制了操作人员的信任度和可审计性。MLLMs具备强大的推理能力,但往往在细粒度空间理解和对象幻觉方面存在不足。此前的工作ByDeWay提出了基于分层深度的提示(LDP),这是一种无需训练的框架,通过利用单目深度估计构建提示来缓解幻觉。然而,粗糙的深度分层无法解决同一几何平面内对象之间的空间关系,比如投影关系(“在左侧”“在上方”)和拓扑关系(“在内部”“接触”)。我们提出了ByDeWay-V2,该框架将显式的空间关系上下文与深度线索相结合,以人类可读的谓词形式呈现,作为下游决策支持的可审计证据。我们使用开放词汇对象检测器YOLO-World-L,计算检测到的对象之间的成对几何关系,并将其作为结构化空间谓词注入到MLLM的提示中,无需任何训练即可弥合3D场景深度与2D空间语义之间的差距。我们在视觉空间推理(VSR)和BLINK基准上对多个MLLMs评估了ByDeWay-V2,通过POPE评估幻觉接地。在BLINK空间子集上,对于Qwen2.5-VL,ByDeWay-V2相比LDP实现了46%的相对F1提升,并且在VSR上将BLIP-Base的空间推理从接近随机的性能提升至具有竞争力的0.53的F1值。我们最轻量化的配置在CPU上严格控制在40个令牌的上下文预算下运行,表明该框架适用于资源受限的实时决策支持场景。

英文摘要

As Multimodal Large Language Models (MLLMs) are increasingly deployed in decision-critical pipelines such as robotics, embodied AI, and safety monitoring, the opacity of their spatial judgments limits operator trust and auditability. MLLMs demonstrate strong reasoning but often struggle with fine-grained spatial understanding and object hallucination. Prior work, ByDeWay, introduced Layered-Depth-Based Prompting (LDP), a training-free framework that mitigates hallucinations by structuring prompts using monocular depth estimation. However, coarse depth layering falls short in resolving object-to-object spatial relationships within the same geometric plane, such as projective ("left of", "above") and topological ("inside", "touching") relations. We propose ByDeWay-V2, which integrates explicit spatial relational context alongside depth cues, expressed as human-readable predicates that serve as auditable evidence for downstream decision support. Using an open-vocabulary object detector (YOLO-World-L), our framework computes pairwise geometric relations between detected objects and injects them as structured spatial predicates into the MLLM prompt, bridging 3D scene depth and 2D spatial semantics without any training. We evaluate ByDeWay-V2 on the Visual Spatial Reasoning (VSR) and BLINK benchmarks across multiple MLLMs, with hallucination grounding assessed via POPE. On the BLINK spatial subset, ByDeWay-V2 achieves a 46 percent relative F1 improvement over LDP for Qwen2.5-VL, and recovers BLIP-Base's spatial reasoning on VSR from near-random performance to a competitive F1 of 0.53. Our lightest configuration operates under a strict 40-token context budget on CPU, showing the framework's suitability for resource-constrained, real-time decision-support settings.

Comments14 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑