arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-12-08 至 2025-12-08 共收录 12 信号源:cs.CV, cs.AI, cs.LG

1. 视觉推理 12 篇

2509.19552 2025-12-08 cs.CV 85%

iFinder: Structured Zero-Shot Vision-Based LLM Grounding for Dash-Cam Video Reasoning

iFinder: 结构化零样本视觉基于LLM的地面定位用于行车记录仪视频推理

Manyi Yao, Bingbing Zhuang, Sparsh Garg, Amit Roy-Chowdhury, Christian Shelton, Manmohan Chandraker, Abhishek Aich

机构 * NEC Laboratories, America(NEC美国实验室) University of California, Riverside(加州大学河滨分校) University of California, San Diego(加州大学圣地亚哥分校)

专题命中 视觉推理 :grounding(title,abstract);vision-language model(abstract);VLM(abstract);分类 cs.CV

AI总结 iFinder通过结构化语义接地框架,利用行车记录仪视频中的关键线索提升LLM在驾驶视频推理中的性能。

Comments Accepted at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16054 2025-12-08 cs.CV 83%

Language-Instructed Reasoning for Group Activity Detection via Multimodal Large Language Model

基于多模态大语言模型的语言引导推理用于群体活动检测

Jihua Peng, Qianxiong Xu, Yichen Liu, Chenxi Liu, Cheng Long, Rui Zhao, Ziyue Li

专题命中 视觉推理 :multimodal large language model(title,abstract);MLLM(abstract);分类 cs.CV

AI总结 本文提出LIR-GAD框架,通过多模态大语言模型实现群体活动检测,引入活动标记和群体标记以提升语义理解和分类性能。

Comments This work is being incorporated into a larger study

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04563 2025-12-08 cs.CV 70%

COOPER: A Unified Model for Cooperative Perception and Reasoning in Spatial Intelligence

COOPER:一种用于空间智能中协作感知与推理的统一模型

Zefeng Zhang, Xiangzhao Hao, Hengzhu Tang, Zhenyu Zhang, Jiawei Sheng, Xiaodong Li, Zhenyang Li, Li Gao, Daiting Shi, Dawei Yin, Tingwen Liu

机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Baidu Inc.(百度公司)

专题命中 视觉推理 :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV

AI总结 COOPER是一种统一的多模态大语言模型,通过整合深度和分割等辅助模态,提升空间感知与推理能力,实现空间智能的增强。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18200 2025-12-08 cs.CV 70%

InfiniBench: Infinite Benchmarking for Visual Spatial Reasoning with Customizable Scene Complexity

InfiniBench:用于可定制场景复杂度的无限基准测试

Haoming Wang, Qiyao Xue, Wei Gao

机构 * University of Pittsburgh(匹兹堡大学)

专题命中 视觉推理 :vision-language model(abstract);VLM(abstract);分类 cs.CV

AI总结 InfiniBench通过可定制的3D场景生成方法,提升视觉语言模型在复杂空间推理任务中的评估能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15436 2025-12-08 cs.CV 70%

Adaptive Chain-of-Focus Reasoning via Dynamic Visual Search and Zooming for Efficient VLMs

基于动态视觉搜索和缩放的自适应聚焦推理方法用于高效VLMs

Xintong Zhang, Zhi Gao, Bofei Zhang, Pengxiang Li, Xiaowen Zhang, Yang Liu, Tao Yuan, Yuwei Wu, Yunde Jia, Song-Chun Zhu, Qing Li

机构 * organization= School of Computer Science \& Technology, Beijing Institute of Technology , city= Beijing , country= China organization= State Key Laboratory of General Artificial Intelligence, BIGAI , city= Beijing , country= China organization= School of Intelligence Science Technology, Peking University , city= Beijing , country= China organization= Guangdong Laboratory of Machine Perception Intelligent Computing, Shenzhen MSU--BIT University , city= Shenzhen , country= China organization= Department of Automation, Tsinghua University , city= Beijing , country= China

专题命中 视觉推理 :vision language model(abstract);visual reasoning(abstract);分类 cs.CV

AI总结 本文提出基于动态视觉搜索和缩放的自适应聚焦推理方法,提升VLMs的多模态推理效率和实际应用效果。

Comments https://github.com/xtong-zhang/Chain-of-Focus

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05809 2025-12-08 cs.CV cs.AI 62%

Probing the effectiveness of World Models for Spatial Reasoning through Test-time Scaling

通过测试时缩放探查世界模型在空间推理中的有效性

Saurav Jha, M. Jehanzeb Mirza, Wei Lin, Shiqi Yang, Sarath Chandar

机构 * MILA – Quebec AI Institute(魁北克人工智能研究所) Polytechnique Montréal(蒙特利尔理工学院) MIT CSAIL(麻省理工学院计算机科学与人工智能实验室) Institute for Machine Learning, Johannes Kepler University Linz(林茨约瑟夫·夫兰克大学机器学习研究所) Nankai University(南开大学)

专题命中 视觉推理 :vision-language model(abstract);分类 cs.CV、cs.AI

AI总结 本文提出ViSA框架,通过可验证的微断言改进世界模型的空间推理能力,但发现当前模型在复杂任务中仍存在信息瓶颈。

Comments Extended abstract at World Modeling Workshop 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16567 2025-12-08 cs.CV cs.AI 62%

V-CECE: Visual Counterfactual Explanations via Conceptual Edits

V-CECE:通过概念编辑生成视觉反事实解释

Nikolaos Spanos, Maria Lymperaiou, Giorgos Filandrianos, Konstantinos Thomas, Athanasios Voulodimos, Giorgos Stamou

机构 * National Technical University of Athens(国家技术大学雅典)

专题命中 视觉推理 :vision language model(abstract);分类 cs.CV、cs.AI

AI总结 V-CECE通过概念编辑生成反事实解释,无需训练即可产生人类水平的可解释性结果。

Comments Accepted in NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.07755 2025-12-08 cs.CV cs.AI cs.GR cs.RO 62%

SAT: Dynamic Spatial Aptitude Training for Multimodal Language Models

SAT:多模态语言模型的动态空间能力训练

Arijit Ray, Jiafei Duan, Ellis Brown, Reuben Tan, Dina Bashkirova, Rose Hendrix, Kiana Ehsani, Aniruddha Kembhavi, Bryan A. Plummer, Ranjay Krishna, Kuo-Hao Zeng, Kate Saenko

机构 * Boston University(波士顿大学) University of Washington(华盛顿大学) Allen Institute for AI(人工智能研究院) Microsoft Research(微软研究院) New York University(纽约大学)

专题命中 视觉推理 :LLaVA(abstract);分类 cs.CV、cs.AI

AI总结 SAT通过模拟数据提升多模态语言模型在动态空间推理中的能力,实验表明其在多个基准测试中优于现有方法。

Comments Accepted to COLM 2025. Project webpage: https://arijitray.com/SAT/

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05965 2025-12-08 cs.CV 57%

EditThinker: Unlocking Iterative Reasoning for Any Image Editor

EditThinker: 解锁任何图像编辑器的迭代推理

Hongyu Li, Manyuan Zhang, Dian Zheng, Ziyu Guo, Yimeng Jia, Kaituo Feng, Hao Yu, Yexin Liu, Yan Feng, Peng Pei, Xunliang Cai, Linjiang Huang, Hongsheng Li, Si Liu

机构 * Beihang University(北航大学) Meituan(美团) CUHK MMLab CUHK IMIXR Tsinghua University(清华大学)

专题命中 视觉推理 :MLLM(abstract);分类 cs.CV

AI总结 EditThinker通过迭代推理框架提升图像编辑指令遵循能力,利用强化学习优化编辑过程,显著提高模型性能。

Comments Project page: https://appletea233.github.io/think-while-edit

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05930 2025-12-08 cs.AI 57%

PRiSM: An Agentic Multimodal Benchmark for Scientific Reasoning via Python-Grounded Evaluation

PRiSM:一种基于Python基础评估的科学推理多模态基准

Shima Imani, Seungwhan Moon, Adel Ahmadyan, Lu Zhang, Kirmani Ahmed, Babak Damavandi

机构 * Meta Reality Lab(Meta现实实验室)

专题命中 视觉推理 :vision-language model(abstract);分类 cs.AI

AI总结 PRiSM通过基于Python代码的动态多模态基准,评估科学推理能力,揭示VLMs在科学领域中的局限性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05398 2025-12-08 cs.CV 57%

The Dynamic Prior: Understanding 3D Structures for Casual Dynamic Videos

动态先验:为随意动态视频理解3D结构

Zhuoyuan Wu, Xurui Yang, Jiahui Huang, Yue Wang, Jun Gao

机构 * PKU(北京大学) NVIDIA(英伟达) USC(南加州大学) University of Michigan(密歇根大学)

专题命中 视觉推理 :vision-language model(abstract);分类 cs.CV

AI总结 本文提出动态先验模型,利用Vision-Language Models和SAM2实现无需任务特定训练的动态物体识别,提升结构3D理解的准确性和鲁棒性。

Comments Code is available at https://github.com/wuzy2115/DYNAPO

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16334 2025-12-08 cs.AI cs.CL 57%

OpenMMReasoner: Pushing the Frontiers for Multimodal Reasoning with an Open and General Recipe

OpenMMReasoner: 推动多模态推理的前沿研究:一个开放且通用的配方

Kaichen Zhang, Keming Wu, Zuhao Yang, Bo Li, Kairui Hu, Bin Wang, Ziwei Liu, Xingxuan Li, Lidong Bing

机构 * MiroMind AI Nanyang Technological University(南洋理工大学) Tsinghua University(清华大学) LMMs-Lab Team(多模态实验室团队)

专题命中 视觉推理 :visual reasoning(abstract);分类 cs.AI

AI总结 OpenMMReasoner提出了一种开放且通用的多模态推理训练配方,通过两阶段方法提升推理性能,实现在多个基准测试中超越现有基线。

详情

展开后加载摘要…

URL PDF HTML 收藏