arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-09-03 至 2025-09-03 共收录 13 信号源:cs.CV, cs.AI, cs.LG

1. 视觉推理 13 篇

2412.14446 2025-09-03 cs.CV cs.LG 88%

VLM-AD: End-to-End Autonomous Driving through Vision-Language Model Supervision

Yi Xu, Yuxin Hu, Zaiwei Zhang, Gregory P. Meyer, Siva Karthik Mustikovela, Siddhartha Srinivasa, Eric M. Wolff, Xin Huang

机构 * Cruise LLC (GM)(Cruise LLC(GM)) Northeastern University(东北大学) OpenAI(开放人工智能研究所) Meta University of Washington(华盛顿大学) Waymo LLC

专题命中 视觉推理 :vision-language model(title,abstract);VLM(title,abstract);分类 cs.CV、cs.LG

Comments Accepted by CoRL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00053 2025-09-03 cs.MM cs.AI cs.CL 88%

Traj-MLLM: Can Multimodal Large Language Models Reform Trajectory Data Mining?

Shuo Liu, Di Yao, Yan Lin, Gao Cong, Jingping Bi

机构 * University of Chinese Academy of Sciences(中国科学院大学) Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所) Department of Computer Science, Aalborg University(奥胡斯大学计算机科学系) College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算与数据科学学院)

专题命中 视觉推理 :multimodal large language model(title,abstract);MLLM(title,abstract);分类 cs.AI

Comments 20 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.09698 2025-09-03 cs.RO cs.AI 79%

ManipBench: Benchmarking Vision-Language Models for Low-Level Robot Manipulation

Enyu Zhao, Vedant Raval, Hejia Zhang, Jiageng Mao, Zeyu Shangguan, Stefanos Nikolaidis, Yue Wang, Daniel Seita

机构 * Department of Computer Science, University of Southern California(计算机科学系,南加州大学)

专题命中 视觉推理 :vision-language model(title,abstract);分类 cs.AI

Comments Conference on Robot Learning (CoRL) 2025. 50 pages and 30 figures. v2 is the camera-ready and includes a few more new experiments compared to v1

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.00711 2025-09-03 cs.CV cs.AI 79%

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework

Chao Wang, Chunbai Zhang, Yongxiao Tian, Yang Zhou, Yan Peng

机构 * School of Future Technology, Shanghai University(未来技术学院,上海大学) School of Mechatronic Engineering and Automation, Shanghai University(机械电子工程与自动化学院,上海大学)

专题命中 视觉推理 :vision-language model(abstract);VLM(abstract);visual reasoning(abstract);分类 cs.CV、cs.AI

Comments 14 pages,17 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.21113 2025-09-03 cs.CV cs.AI cs.LG 75%

R-4B: Incentivizing General-Purpose Auto-Thinking Capability in MLLMs via Bi-Mode Annealing and Reinforce Learning

Qi Yang, Bolin Ni, Shiming Xiang, Han Hu, Houwen Peng, Jie Jiang

机构 * Tencent Hunyuan Team(腾讯文言团队) Institute of Automation, CAS(中国科学院自动化研究所)

专题命中 视觉推理 :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI、cs.LG

Comments 20 pages, 14 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.16044 2025-09-03 cs.CV 70%

ZoomEye: Enhancing Multimodal LLMs with Human-Like Zooming Capabilities through Tree-Based Image Exploration

Haozhan Shen, Kangjia Zhao, Tiancheng Zhao, Ruochen Xu, Zilun Zhang, Mingwei Zhu, Jianwei Yin

机构 * Zhejiang University(浙江大学) Om AI Research(Om AI 研究所) Binjiang Institute of Zhejiang University(浙江大学滨江研究院)

专题命中 视觉推理 :visual reasoning(abstract);multimodal large language model(abstract);分类 cs.CV

Comments Accepted by EMNLP-2025 Main. Project page: https://szhanz.github.io/zoomeye/

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.18798 2025-09-03 cs.CL 67%

Distill Visual Chart Reasoning Ability from LLMs to MLLMs

Wei He, Zhiheng Xi, Wanxu Zhao, Xiaoran Fan, Yiwen Ding, Zifei Shan, Tao Gui, Qi Zhang, Xuanjing Huang

机构 * School of Computer Science, Fudan University(复旦大学计算机学院) Weixin Group, Tencent(腾讯Weixin部门) Shanghai Innovation Institute(上海创新研究院) Shanghai Key Lab of Intelligent Information Processing(上海智能信息处理重点实验室)

专题命中 视觉推理 :visual reasoning(abstract);multimodal large language model(abstract)

Comments Accepted to EMNLP 2025 Findings. The code and dataset are publicly available at https://github.com/hewei2001/ReachQA

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00484 2025-09-03 cs.CV cs.AI 62%

VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding

Zhihong Zhang, Xiaojian Huang, Jin Xu, Zhuodong Luo, Xinzhi Wang, Jiansheng Wei, Xuejin Chen

机构 * University of Science and Technology of China(中国科学技术大学) Huawei Noah’s Ark Lab(华为诺亚实验室)

专题命中 视觉推理 :vision language model(abstract);分类 cs.CV、cs.AI

Comments https://videorewardbench.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00465 2025-09-03 cs.RO cs.AI cs.CV 62%

Embodied Spatial Intelligence: from Implicit Scene Modeling to Spatial Reasoning

Jiading Fang

机构 * Toyota Technological Institute at Chicago (TTIC)(丰田技术研究所(芝加哥))

专题命中 视觉推理 :grounding(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20356 2025-09-03 cs.CV 57%

Detecting Visual Information Manipulation Attacks in Augmented Reality: A Multimodal Semantic Reasoning Approach

Yanming Xiu, Maria Gorlatova

机构 * Duke University(杜克大学)

专题命中 视觉推理 :vision-language model(abstract);分类 cs.CV

Comments The paper has been accepted to the 2025 IEEE International Symposium on Mixed and Augmented Reality (ISMAR), and selected for publication in the 2025 IEEE Transactions on Visualization and Computer Graphics (TVCG) special issue

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01656 2025-09-03 cs.CV cs.CL 57%

Reinforced Visual Perception with Tools

Zetong Zhou, Dongping Chen, Zixian Ma, Zhihan Hu, Mingyang Fu, Sinan Wang, Yao Wan, Zhou Zhao, Ranjay Krishna

机构 * ONE Lab, HUST(华中科技大学 ONE 实验室) ONE Lab, HUST University of Maryland(华中科技大学 与 马里兰大学 ONE 实验室) University of Washington(华盛顿大学) Zhejiang University(浙江大学)

专题命中 视觉推理 :visual reasoning(abstract);分类 cs.CV

Comments Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01209 2025-09-03 cs.CV 57%

Measuring Image-Relation Alignment: Reference-Free Evaluation of VLMs and Synthetic Pre-training for Open-Vocabulary Scene Graph Generation

Maëlic Neau, Zoe Falomir, Cédric Buche, Akihiro Sugimoto

机构 * Computing Science Department, Umeå University(乌梅大学计算科学系) CNRS IRL 2010 CROSSING(法国CNRS IRL 2010 CROSSING) IMT Atlantique(IMT阿蒂昂大学) National Institute of Informatics(日本信息机构)

专题命中 视觉推理 :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.06661 2025-09-03 cs.RO 50%

Domain-Conditioned Scene Graphs for State-Grounded Task Planning

Jonas Herzog, Jiangpin Liu, Yue Wang

机构 * Zhejiang University(浙江大学)

专题命中 视觉推理 :grounding(abstract)

Comments Accepted for IROS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏