发表机构
The University of Hong Kong; University of Southern California; China Telecom; Columbia University; City University of Hong Kong; Chinese Academy of Sciences(香港大学; 南加州大学; 中国电信; 哥伦比亚大学; 香港城市大学; 中国科学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
INTCORT通过输入变换和置信度路由,在不修改模型内部机制的情况下,显著提升视觉-语言模型的空间推理能力,平均提升10.01%。
AI 中文摘要
视觉-语言模型(VLMs)在多模态任务中展现了卓越的能力,但它们在空间推理方面的表现仍然不佳。现有的依赖训练和无训练增强方法分别存在高计算成本并导致灾难性遗忘,以及内部机制干扰而损害通用能力的问题。在这项工作中,我们首先验证了两个关键假设:适当的几何图像变换和查询反转变换可以纠正错误的空间预测,并且正确的预测在关系令牌上表现出比错误预测更高的置信度。基于这些发现,我们提出了INTCORT,一个无训练的空间推理增强框架,通过输入变换构建多个推理视图,并通过关系令牌置信度路由聚合它们的预测,而无需修改VLM的内部机制。在多个常用基准上的实验结果表明,INTCORT在多种VLM上显著提高了空间推理准确性,在所有模型和基准上平均提升了10.01%。与先前工作相比,INTCORT实现了更优的性能,提升幅度高达25.01%。
英文摘要
Vision-Language Models (VLMs) have demonstrated remarkable capabilities in multimodal tasks, yet they still exhibit poor ability in spatial reasoning. Existing training-dependent and training-free enhancement methods suffer from high computational costs with catastrophic forgetting and internal mechanism interference that compromises general capabilities, respectively. In this work, we first verify two key hypotheses: appropriate geometric image transformation and query-reversal transformation can recover incorrect spatial predictions, and correct predictions exhibit higher relation-token confidence than incorrect ones. Based on these findings, we propose INTCORT, a training-free spatial reasoning enhancement framework that constructs multiple inference views through input transformations and aggregates their predictions via relation-token confidence routing, without modifying the VLM's internal mechanisms. Experimental results on several commonly-used benchmarks demonstrate that INTCORT substantially improves spatial reasoning accuracy across diverse VLMs, achieving an average improvement of 10.01% over all models and benchmarks. Compared with prior works, INTCORT achieves superior performance with improvements of up to 25.01%.