arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-09-23 至 2025-09-23 共收录 12 信号源:cs.CV, cs.AI, cs.LG

1. 视觉推理 12 篇

2509.16517 2025-09-23 cs.CV cs.AI cs.CL cs.MM 91%

Seeing Culture: A Benchmark for Visual Reasoning and Grounding

Burak Satar, Zhixin Ma, Patrick A. Irawan, Wilfried A. Mulyawan, Jing Jiang, Ee-Peng Lim, Chong-Wah Ngo

机构 * Singapore Management University(新加坡管理大学) Bandung Institute of Technology(班达理工大学)

专题命中 视觉推理 :visual reasoning(title,abstract);grounding(title,abstract);vision-language model(abstract);visual question answering(abstract)

Comments Accepted to EMNLP 2025 Main Conference, https://seeingculture-benchmark.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17042 2025-09-23 cs.RO 82%

Orchestrate, Generate, Reflect: A VLM-Based Multi-Agent Collaboration Framework for Automated Driving Policy Learning

Zengqi Peng, Yusen Xie, Yubin Wang, Rui Yang, Qifeng Chen, Jun Ma

机构 * Robotics and Autonomous Systems Thrust, The Hong Kong University of Science and Technology (Guangzhou)(机器人与自主系统方向,香港科学与技术大学(广州)) Department of Computer Science and Engineering, The Hong Kong University of Science and Technology(计算机科学与工程系,香港科学与技术大学) Cheng Kar-Shun Robotics Institute, The Hong Kong University of Science and Technology(陈家骏机器人研究所,香港科学与技术大学)

专题命中 视觉推理 :VLM(title,abstract);vision-language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05255 2025-09-23 cs.CV cs.CL 79%

Open Vision Reasoner: Transferring Linguistic Cognitive Behavior for Visual Reasoning

Yana Wei, Liang Zhao, Jianjian Sun, Kangheng Lin, Jisheng Yin, Jingcheng Hu, Yinmin Zhang, En Yu, Haoran Lv, Zejia Weng, Jia Wang, Chunrui Han, Yuang Peng, Qi Han, Zheng Ge, Xiangyu Zhang, Daxin Jiang, Vishal M. Patel

机构 * Johns Hopkins University(约翰霍普金斯大学) StepFun BUPT(北京邮电大学) UCAS(中国科学院大学) THU(清华大学) HUST(华中科技大学)

专题命中 视觉推理 :visual reasoning(title,abstract);分类 cs.CV

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20024 2025-09-23 cs.CV cs.AI cs.RO 73%

ReasonPlan: Unified Scene Prediction and Decision Reasoning for Closed-loop Autonomous Driving

Xueyi Liu, Zuodong Zhong, Yuxin Guo, Yun-Fu Liu, Zhiguo Su, Qichao Zhang, Junli Wang, Yinfeng Gao, Yupeng Zheng, Qiao Lin, Huiyong Chen, Dongbin Zhao

机构 * SKL-MAIS, Institute of Automation, Chinese Academy of Sciences, Beijing, China(中国科学院自动化研究所SKL-MAIS部门,北京) School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China(中国科学院大学人工智能学院,北京) EACON, Fujian, China(福建EACON机构,中国) School of Automation and Electrical Engineering, University of Science and Technology Beijing, Beijing, China(北京科技大学自动化与电气工程学院)

专题命中 视觉推理 :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments 18 pages; 9 figures; https://github.com/Liuxueyi/ReasonPlan

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16597 2025-09-23 cs.CL 71%

MCP: A Control-Theoretic Orchestration Framework for Synergistic Efficiency and Interpretability in Multimodal Large Language Models

Luyan Zhang

机构 * Northeastern University(东北大学)

专题命中 视觉推理 :multimodal large language model(title)

Comments 13 pages, 6 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17437 2025-09-23 cs.CL 67%

GeoPQA: Bridging the Visual Perception Gap in MLLMs for Geometric Reasoning

Guizhen Chen, Weiwen Xu, Hao Zhang, Hou Pong Chan, Deli Zhao, Anh Tuan Luu, Yu Rong

机构 * Nanyang Technological University(南洋理工大学) DAMO Academy, Alibaba Group(阿里达摩院) Hupan Lab(华普实验室)

专题命中 视觉推理 :grounding(abstract);MLLM(abstract)

Comments Accepted to EMNLP2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.16989 2025-09-23 cs.CL 67%

All-in-one: Understanding and Generation in Multimodal Reasoning with the MAIA Benchmark

Davide Testa, Giovanni Bonetta, Raffaella Bernardi, Alessandro Bondielli, Alessandro Lenci, Alessio Miaschi, Lucia Passaro, Bernardo Magnini

机构 * Università di Roma La Sapienza(罗马La Sapienza大学) Fondazione Bruno Kessler (FBK)(布鲁诺·凯斯勒基金会) Free University of Bozen-Bolzano(博兹纳-博尔扎诺自由大学) Dept. of Computer Science, University of Pisa(比萨大学计算机科学系) CoLing Lab, Dept. of Philology, Literature and Linguistics, University of Pisa(比萨大学语言学、文学与语言学系CoLing实验室) Istituto di Linguistica Computazionale "A. Zampolli" (CNR-ILC), ItaliaNLP Lab, Pisa(A. Zampolli计算语言学研究所(CNR-ILC),意大利NLP实验室,比萨)

专题命中 视觉推理 :vision language model(abstract);visual language model(abstract)

Comments Accepted at Findings of EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16721 2025-09-23 cs.CV cs.AI cs.RO 62%

Text-Scene: A Scene-to-Language Parsing Framework for 3D Scene Understanding

Haoyuan Li, Rui Liu, Hehe Fan, Yi Yang

机构 * College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院)

专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.CV、cs.AI

Comments 19 pages, 12 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15293 2025-09-23 cs.CV cs.RO 57%

How Good are Foundation Models in Step-by-Step Embodied Reasoning?

Dinura Dissanayake, Ahmed Heakl, Omkar Thawakar, Noor Ahsan, Ritesh Thawkar, Ketan More, Jean Lahoud, Rao Anwer, Hisham Cholakkal, Ivan Laptev, Fahad Shahbaz Khan, Salman Khan

机构 * Mohamed bin Zayed University of AI(穆罕默德·本·扎耶德人工智能大学) Linköping University(林雪平大学) Australian National University(澳大利亚国立大学)

专题命中 视觉推理 :grounding(abstract);分类 cs.CV

Comments Project page: https://mbzuai-oryx.github.io/FoMER-Bench/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17740 2025-09-23 cs.CV cs.CL 57%

WISE: Weak-Supervision-Guided Step-by-Step Explanations for Multimodal LLMs in Image Classification

Yiwen Jiang, Deval Mehta, Siyuan Yan, Yaling Shen, Zimu Wang, Zongyuan Ge

机构 * Faculty of Engineering, Monash University(墨尔本大学工程学院) AIM for Health Lab, Faculty of IT, Monash University(墨尔本大学信息技术学院健康人工智能实验室)

专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.CV

Comments Accepted at EMNLP 2025 (Main)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17044 2025-09-23 cs.CV 57%

AgriDoctor: A Multimodal Intelligent Assistant for Agriculture

Mingqing Zhang, Zhuoning Xu, Peijie Wang, Rongji Li, Liang Wang, Qiang Liu, Jian Xu, Xuyao Zhang, Shu Wu, Liang Wang

机构 * State Key Laboratory of Multimodal Artificial Intelligence Systems(多模态人工智能系统国家重点实验室) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)

专题命中 视觉推理 :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16226 2025-09-23 cs.CL cs.AI 57%

On LLM-Based Scientific Inductive Reasoning Beyond Equations

Brian S. Lin, Jiaxin Yuan, Zihan Zhou, Shouli Wang, Shuo Wang, Cunliang Kong, Qi Shi, Yuxuan Li, Liner Yang, Zhiyuan Liu, Maosong Sun

机构 * Dept. of Comp. Sci. & Tech., Institute for AI, BNRist Center, Tsinghua University(计算机科学与技术系,人工智能研究院,BNRist中心,清华大学) Jiangsu Collaborative Innovation Center for Language Ability, Jiangsu Normal University(江苏语言能力协同创新中心,江苏师范大学) Beijing Language and Culture University(北京语言文化大学) Xiamen University(厦门大学) Harbin Institute of Technology(哈尔滨工业大学)

专题命中 视觉推理 :grounding(abstract);分类 cs.AI

Comments 24 pages

详情

展开后加载摘要…

URL PDF HTML 收藏