arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-09-03 至 2025-09-03 共收录 72 信号源:cs.CV, cs.AI, cs.LG

1. 视觉问答 4 篇

2507.22369 2025-09-03 cs.CV cs.AI 81%

Exploring the Application of Visual Question Answering (VQA) for Classroom Activity Monitoring

Sinh Trong Vu, Hieu Trung Pham, Dung Manh Nguyen, Hieu Minh Hoang, Nhu Hoang Le, Thu Ha Pham, Tai Tan Mai

机构 * Banking Academy of Vietnam(越南银行学院) Dublin City University(都柏林城市大学)

专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19944 2025-09-03 cs.CV cs.CL 77%

KRETA: A Benchmark for Korean Reading and Reasoning in Text-Rich VQA Attuned to Diverse Visual Contexts

Taebaek Hwang, Minseo Kim, Gisang Lee, Seonuk Kim, Hyunjun Eun

机构 * Waddle Seoul National University(首尔国立大学) Krafton UNIST(全南大学) SK Telecom(SK电信)

专题命中 视觉问答 :vision-language model(abstract);VLM(abstract);visual question answering(abstract);分类 cs.CV

Comments Accepted to EMNLP 2025 (Main Conference)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02164 2025-09-03 cs.CV 70%

Omnidirectional Spatial Modeling from Correlated Panoramas

Xinshen Zhang, Tongxi Fu, Xu Zheng

机构 * The Hong Kong Polytechnic University(香港理工大学) Zhejiang Sci-Tech University(浙江科技学院)

专题命中 视觉问答 :visual question answering(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.02748 2025-09-03 cs.CV cs.AI 62%

Story Generation from Visual Inputs: Techniques, Related Tasks, and Challenges

Daniel A. P. Oliveira, Eugénio Ribeiro, David Martins de Matos

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.AI

Comments 23 pages

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 视觉推理 13 篇

2412.14446 2025-09-03 cs.CV cs.LG 88%

VLM-AD: End-to-End Autonomous Driving through Vision-Language Model Supervision

Yi Xu, Yuxin Hu, Zaiwei Zhang, Gregory P. Meyer, Siva Karthik Mustikovela, Siddhartha Srinivasa, Eric M. Wolff, Xin Huang

机构 * Cruise LLC (GM)(Cruise LLC(GM)) Northeastern University(东北大学) OpenAI(开放人工智能研究所) Meta University of Washington(华盛顿大学) Waymo LLC

专题命中 视觉推理 :vision-language model(title,abstract);VLM(title,abstract);分类 cs.CV、cs.LG

Comments Accepted by CoRL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00053 2025-09-03 cs.MM cs.AI cs.CL 88%

Traj-MLLM: Can Multimodal Large Language Models Reform Trajectory Data Mining?

Shuo Liu, Di Yao, Yan Lin, Gao Cong, Jingping Bi

机构 * University of Chinese Academy of Sciences(中国科学院大学) Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所) Department of Computer Science, Aalborg University(奥胡斯大学计算机科学系) College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算与数据科学学院)

专题命中 视觉推理 :multimodal large language model(title,abstract);MLLM(title,abstract);分类 cs.AI

Comments 20 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.09698 2025-09-03 cs.RO cs.AI 79%

ManipBench: Benchmarking Vision-Language Models for Low-Level Robot Manipulation

Enyu Zhao, Vedant Raval, Hejia Zhang, Jiageng Mao, Zeyu Shangguan, Stefanos Nikolaidis, Yue Wang, Daniel Seita

机构 * Department of Computer Science, University of Southern California(计算机科学系,南加州大学)

专题命中 视觉推理 :vision-language model(title,abstract);分类 cs.AI

Comments Conference on Robot Learning (CoRL) 2025. 50 pages and 30 figures. v2 is the camera-ready and includes a few more new experiments compared to v1

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.00711 2025-09-03 cs.CV cs.AI 79%

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework

Chao Wang, Chunbai Zhang, Yongxiao Tian, Yang Zhou, Yan Peng

机构 * School of Future Technology, Shanghai University(未来技术学院,上海大学) School of Mechatronic Engineering and Automation, Shanghai University(机械电子工程与自动化学院,上海大学)

专题命中 视觉推理 :vision-language model(abstract);VLM(abstract);visual reasoning(abstract);分类 cs.CV、cs.AI

Comments 14 pages,17 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.21113 2025-09-03 cs.CV cs.AI cs.LG 75%

R-4B: Incentivizing General-Purpose Auto-Thinking Capability in MLLMs via Bi-Mode Annealing and Reinforce Learning

Qi Yang, Bolin Ni, Shiming Xiang, Han Hu, Houwen Peng, Jie Jiang

机构 * Tencent Hunyuan Team(腾讯文言团队) Institute of Automation, CAS(中国科学院自动化研究所)

专题命中 视觉推理 :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI、cs.LG

Comments 20 pages, 14 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.16044 2025-09-03 cs.CV 70%

ZoomEye: Enhancing Multimodal LLMs with Human-Like Zooming Capabilities through Tree-Based Image Exploration

Haozhan Shen, Kangjia Zhao, Tiancheng Zhao, Ruochen Xu, Zilun Zhang, Mingwei Zhu, Jianwei Yin

机构 * Zhejiang University(浙江大学) Om AI Research(Om AI 研究所) Binjiang Institute of Zhejiang University(浙江大学滨江研究院)

专题命中 视觉推理 :visual reasoning(abstract);multimodal large language model(abstract);分类 cs.CV

Comments Accepted by EMNLP-2025 Main. Project page: https://szhanz.github.io/zoomeye/

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.18798 2025-09-03 cs.CL 67%

Distill Visual Chart Reasoning Ability from LLMs to MLLMs

Wei He, Zhiheng Xi, Wanxu Zhao, Xiaoran Fan, Yiwen Ding, Zifei Shan, Tao Gui, Qi Zhang, Xuanjing Huang

机构 * School of Computer Science, Fudan University(复旦大学计算机学院) Weixin Group, Tencent(腾讯Weixin部门) Shanghai Innovation Institute(上海创新研究院) Shanghai Key Lab of Intelligent Information Processing(上海智能信息处理重点实验室)

专题命中 视觉推理 :visual reasoning(abstract);multimodal large language model(abstract)

Comments Accepted to EMNLP 2025 Findings. The code and dataset are publicly available at https://github.com/hewei2001/ReachQA

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00484 2025-09-03 cs.CV cs.AI 62%

VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding

Zhihong Zhang, Xiaojian Huang, Jin Xu, Zhuodong Luo, Xinzhi Wang, Jiansheng Wei, Xuejin Chen

机构 * University of Science and Technology of China(中国科学技术大学) Huawei Noah’s Ark Lab(华为诺亚实验室)

专题命中 视觉推理 :vision language model(abstract);分类 cs.CV、cs.AI

Comments https://videorewardbench.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00465 2025-09-03 cs.RO cs.AI cs.CV 62%

Embodied Spatial Intelligence: from Implicit Scene Modeling to Spatial Reasoning

Jiading Fang

机构 * Toyota Technological Institute at Chicago (TTIC)(丰田技术研究所(芝加哥))

专题命中 视觉推理 :grounding(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20356 2025-09-03 cs.CV 57%

Detecting Visual Information Manipulation Attacks in Augmented Reality: A Multimodal Semantic Reasoning Approach

Yanming Xiu, Maria Gorlatova

机构 * Duke University(杜克大学)

专题命中 视觉推理 :vision-language model(abstract);分类 cs.CV

Comments The paper has been accepted to the 2025 IEEE International Symposium on Mixed and Augmented Reality (ISMAR), and selected for publication in the 2025 IEEE Transactions on Visualization and Computer Graphics (TVCG) special issue

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01656 2025-09-03 cs.CV cs.CL 57%

Reinforced Visual Perception with Tools

Zetong Zhou, Dongping Chen, Zixian Ma, Zhihan Hu, Mingyang Fu, Sinan Wang, Yao Wan, Zhou Zhao, Ranjay Krishna

机构 * ONE Lab, HUST(华中科技大学 ONE 实验室) ONE Lab, HUST University of Maryland(华中科技大学 与 马里兰大学 ONE 实验室) University of Washington(华盛顿大学) Zhejiang University(浙江大学)

专题命中 视觉推理 :visual reasoning(abstract);分类 cs.CV

Comments Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01209 2025-09-03 cs.CV 57%

Measuring Image-Relation Alignment: Reference-Free Evaluation of VLMs and Synthetic Pre-training for Open-Vocabulary Scene Graph Generation

Maëlic Neau, Zoe Falomir, Cédric Buche, Akihiro Sugimoto

机构 * Computing Science Department, Umeå University(乌梅大学计算科学系) CNRS IRL 2010 CROSSING(法国CNRS IRL 2010 CROSSING) IMT Atlantique(IMT阿蒂昂大学) National Institute of Informatics(日本信息机构)

专题命中 视觉推理 :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.06661 2025-09-03 cs.RO 50%

Domain-Conditioned Scene Graphs for State-Grounded Task Planning

Jonas Herzog, Jiangpin Liu, Yue Wang

机构 * Zhejiang University(浙江大学)

专题命中 视觉推理 :grounding(abstract)

Comments Accepted for IROS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视觉定位与Grounding 16 篇

2509.01554 2025-09-03 cs.CV cs.AI cs.LG 85%

Unified Supervision For Vision-Language Modeling in 3D Computed Tomography

Hao-Chih Lee, Zelong Liu, Hamza Ahmed, Spencer Kim, Sean Huver, Vishwesh Nath, Zahi A. Fayad, Timothy Deyer, Xueyan Mei

专题命中 视觉定位与Grounding :vision-language model(title,abstract);VLM(abstract,comments);分类 cs.CV、cs.AI、cs.LG

Comments ICCV 2025 VLM 3d Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16647 2025-09-03 cs.CV cs.AI 81%

Point, Detect, Count: Multi-Task Medical Image Understanding with Instruction-Tuned Vision-Language Models

Sushant Gautam, Michael A. Riegler, Pål Halvorsen

机构 * Simula Metropolitan Center for Digital Engineering (SimulaMet), Norway(Simula数字工程中心(SimulaMet)) Oslo Metropolitan University (OsloMet), Norway(奥斯陆 Metropolitan 大学(OsloMet)) Simula Research Laboratory, Norway(Simula研究实验室)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV、cs.AI

Comments Accepted as a full paper at the 38th IEEE International Symposium on Computer-Based Medical Systems (CBMS) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02324 2025-09-03 cs.RO 75%

Language-Guided Long Horizon Manipulation with LLM-based Planning and Visual Perception

Changshi Zhou, Haichuan Xu, Ningquan Gu, Zhipeng Wang, Bin Cheng, Pengpeng Zhang, Yanchao Dong, Mitsuhiro Hayashibe, Yanmin Zhou, Bin He

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.13919 2025-09-03 cs.CV cs.AI cs.CL cs.LG cs.RO 75%

Temporal Preference Optimization for Long-Form Video Understanding

Rui Li, Xiaohan Wang, Yuhui Zhang, Orr Zohar, Zeyu Wang, Serena Yeung-Levy

机构 * Stanford University(斯坦福大学)

专题命中 视觉定位与Grounding :LLaVA(abstract);grounding(abstract);分类 cs.CV、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00284 2025-09-03 cs.CV cs.AI 73%

Generative AI for Industrial Contour Detection: A Language-Guided Vision System

Liang Gong, Tommy, Wang, Sara Chaker, Yanchen Dong, Fouad Bousetouane, Brenden Morton, Mark Mendez

机构 * The University of Chicago(芝加哥大学) FabTrack

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.CV、cs.AI

Comments 20 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.16680 2025-09-03 cs.CV 70%

AeroReformer: Aerial Referring Transformer for UAV-based Referring Image Segmentation

Rui Li, Xiaowei Zhao

机构 * Intelligent Control \& Smart Energy (ICSE) Research Group, School of Engineering, University of Warwick, Coventry, CV4 7AL, UK

专题命中 视觉定位与Grounding :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04549 2025-09-03 cs.CV cs.AI cs.MM 62%

MSC: A Marine Wildlife Video Dataset with Grounded Segmentation and Clip-Level Captioning

Quang-Trung Truong, Yuk-Kwan Wong, Vo Hoang Kim Tuyen Dang, Rinaldi Gotama, Duc Thanh Nguyen, Sai-Kit Yeung

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学) Ho Chi Minh University of Science(胡志明市科学大学) Deakin University(德肯大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI

Comments Published at ACMMM2025 (Dataset track)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01716 2025-09-03 cs.AI cs.CL 57%

An LLM-enabled semantic-centric framework to consume privacy policies

Rui Zhao, Vladyslav Melnychuk, Jun Zhao, Jesse Wright, Nigel Shadbolt

机构 * University of Oxford(牛津大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.21484 2025-09-03 q-bio.QM cs.LG stat.ML 57%

Data-driven Discovery of Digital Twins in Biomedical Research

Clémence Métayer, Annabelle Ballesta, Julien Martinelli

机构 * Inserm U1331, Institut Curie Saint-Cloud, France(法国国家医学研究院U1331,圣克鲁医院) Aalto University Espoo, Finland(芬兰艾尔托大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10292 2025-09-03 cs.CV cs.CL 57%

StoryReasoning Dataset: Using Chain-of-Thought for Scene Understanding and Grounded Story Generation

Daniel A. P. Oliveira, David Martins de Matos

机构 * Instituto Superior Técnico, Universidade de Lisboa(里斯本大学技术学院) INESC-ID Lisboa(里斯本INESC-ID)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments 31 pages, 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.00600 2025-09-03 cs.DB cs.AI cs.CL 57%

Semantic Integrity Constraints: Declarative Guardrails for AI-Augmented Data Processing Systems

Alexander W. Lee, Justin Chan, Michael Fu, Nicolas Kim, Akshay Mehta, Deepti Raghavan, Ugur Cetintemel

机构 * Brown University(布朗大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Journal ref PVLDB, 18(11): 4073 - 4080, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.07268 2025-09-03 cs.MM cs.CL cs.CV 57%

Advancing Grounded Multimodal Named Entity Recognition via LLM-Based Reformulation and Box-Based Segmentation

Jinyuan Li, Ziyan Li, Han Li, Jianfei Yu, Rui Xia, Di Sun, Gang Pan

机构 * College of Intelligence and Computing, Tianjin University(智能与计算学院,天津大学) NJUST(南京理工大学) College of Mathematics, Taiyuan University of Technology(数学学院,太原科技大学) Tianjin University of Science and Technology(天津科技大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments Extension of our Findings of EMNLP 2023 & ACL 2024 paper, IEEE Transactions on Multimedia accepted on July 19, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02425 2025-09-03 cs.RO cs.HC 50%

OpenGuide: Assistive Object Retrieval in Indoor Spaces for Individuals with Visual Impairments

Yifan Xu, Qianwei Wang, Vineet Kamat, Carol Menassa

机构 * Department of Civil and Environmental Engineering, University of Michigan(密歇根大学土木与环境工程系) College of Literature, Science, and the Arts, University of Michigan(密歇根大学文学、科学与艺术学院)

专题命中 视觉定位与Grounding :VLM(abstract)

Comments 32 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏