arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-08-08 至 2025-08-08 共收录 38 信号源:cs.CV, cs.AI, cs.LG

1. 视觉问答 4 篇

2506.21586 2025-08-08 cs.CL cs.AI cs.CV 81%

Can Vision Language Models Understand Mimed Actions?

Hyundong Cho, Spencer Lin, Tejas Srinivasan, Michael Saxon, Deuksin Kwon, Natali T. Chavez, Jonathan May

机构 * Information Sciences Institute(信息科学研究所) Institute for Creative Technologies(创意技术研究所) Department of Computer Science(计算机科学系) University of Southern California(南加州大学) University of California, Santa Barbara(加州大学圣巴巴拉分校) Aristotle University of Thessaloniki(希腊雅典纳大学)

专题命中 视觉问答 :vision language model(title);vision-language model(abstract);分类 cs.CV、cs.AI

Comments ACL 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.18351 2025-08-08 cs.CL cs.AI 79%

Multi-Agents Based on Large Language Models for Knowledge-based Visual Question Answering

Zhongjian Hu, Peng Yang, Bing Li, Zhenqi Wang

专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.AI

Comments We would like to withdraw this submission due to ongoing internal review and coordination among the author team. Upon the supervisor's recommendation, we have decided to delay public dissemination until the manuscript undergoes further refinement and aligns with our intended academic trajectory

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.16936 2025-08-08 cs.CL cs.AI 79%

Rationale-guided Prompting for Knowledge-based Visual Question Answering

Zhongjian Hu, Peng Yang, Bing Li, Fengyuan Liu

专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.AI

Comments We would like to withdraw this submission due to ongoing internal review and coordination among the author team. Upon the supervisor's recommendation, we have decided to delay public dissemination until the manuscript undergoes further refinement and aligns with our intended academic trajectory

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04895 2025-08-08 cs.SE 71%

Automated Bug Frame Retrieval from Gameplay Videos Using Vision-Language Models

Wentao Lu, Alexander Senchenko, Abram Hindle, Cor-Paul Bezemer

专题命中 视觉问答 :vision-language model(title)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 视觉推理 10 篇

2508.02095 2025-08-08 cs.CV cs.AI 86%

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models

Shijie Zhou, Alexander Vilesov, Xuehai He, Ziyu Wan, Shuwang Zhang, Aditya Nagachandra, Di Chang, Dongdong Chen, Xin Eric Wang, Achuta Kadambi

机构 * UCLA(美国大学洛杉矶分校) Microsoft(微软公司) UCSC(加州大学圣塔克拉拉分校) USC(美国大学洛杉矶分校)

专题命中 视觉推理 :vision language model(title,abstract);visual reasoning(abstract);grounding(abstract);分类 cs.CV、cs.AI

Comments ICCV 2025, Project Website: https://vlm4d.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21745 2025-08-08 cs.CV 81%

Few-Shot Vision-Language Reasoning for Satellite Imagery via Verifiable Rewards

Aybora Koksal, A. Aydin Alatan

专题命中 视觉推理 :vision-language model(abstract);VLM(abstract);visual question answering(abstract);grounding(abstract)

Comments ICCV 2025 Workshop on Curated Data for Efficient Learning (CDEL). 10 pages, 3 figures, 6 tables. Our model, training code and dataset will be at https://github.com/aybora/FewShotReasoning

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04931 2025-08-08 cs.RO cs.AI 79%

INTENTION: Inferring Tendencies of Humanoid Robot Motion Through Interactive Intuition and Grounded VLM

Jin Wang, Weijie Wang, Boyuan Deng, Heng Zhang, Rui Dai, Nikos Tsagarakis

专题命中 视觉推理 :VLM(title);vision-language model(abstract);分类 cs.AI

Comments Project Web: https://robo-intention.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20766 2025-08-08 cs.CV 77%

Learning Only with Images: Visual Reinforcement Learning with Reasoning, Rendering, and Visual Feedback

Yang Chen, Yufan Shen, Wenxuan Huang, Sheng Zhou, Qunshu Lin, Xinyu Cai, Zhi Yu, Jiajun Bu, Botian Shi, Yu Qiao

专题命中 视觉推理 :visual reasoning(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05405 2025-08-08 cs.AI 70%

DeepPHY: Benchmarking Agentic VLMs on Physical Reasoning

Xinrun Xu, Pi Bu, Ye Wang, Börje F. Karlsson, Ziming Wang, Tengtao Song, Qi Zhu, Jun Song, Zhiming Ding, Bo Zheng

机构 * 1 Taobao \& Tmall Group of Alibaba, 2 Institute of Software, Chinese Academy of Science, 3 University of Chinese Academy of Sciences, 4 Renmin University of China, 5 Informatics Department, PUC-Rio

专题命中 视觉推理 :vision language model(abstract);visual reasoning(abstract);分类 cs.AI

Comments 48 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05234 2025-08-08 cs.CL cs.AI 70%

Resource-Limited Joint Multimodal Sentiment Reasoning and Classification via Chain-of-Thought Enhancement and Distillation

Haonan Shangguan, Xiaocui Yang, Shi Feng, Daling Wang, Yifei Zhang, Ge Yu

专题命中 视觉推理 :multimodal large language model(abstract);MLLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05221 2025-08-08 cs.CV cs.AI cs.LG 67%

ReasoningTrack: Chain-of-Thought Reasoning for Long-term Vision-Language Tracking

Xiao Wang, Liye Jin, Xufeng Lou, Shiao Wang, Lan Chen, Bo Jiang, Zhipeng Zhang

机构 * School of Computer Science and Technology, Anhui University(安徽大学计算机科学与技术学院) School of Artificial Intelligence, Shanghai Jiao Tong University(上海交通大学人工智能学院) School of Electronic and Information Engineering, Anhui University(安徽大学电子与信息工程学院)

专题命中 视觉推理 :vision-language model(abstract);分类 cs.CV、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04848 2025-08-08 cs.AI 61%

Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning

Chang Tian, Matthew B. Blaschko, Mingzhe Xing, Xiuxing Li, Yinliang Yue, Marie-Francine Moens

专题命中 视觉推理 :vision-language model(abstract,comments);分类 cs.AI

Comments large language models, large vision-language model, reasoning, non-ideal conditions, reinforcement learning

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05383 2025-08-08 cs.AI 57%

StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models

Xiangxiang Zhang, Jingxuan Wei, Donghong Zhong, Qi Chen, Caijun Jia, Cheng Tan, Jinming Gu, Xiaobo Qin, Zhiping Liu, Liang Hu, Tong Sun, Yuchen Wu, Zewei Sun, Chenwei Lou, Hua Zheng, Tianyang Zhan, Changbao Wang, Shuangzhi Wu, Zefa Lin, Chang Guo, Sihang Yuan, Riwei Chen, Shixiong Zhao, Yingping Zhang, Gaowei Wu, Bihui Yu, Jiahui Wu, Zhehui Zhao, Qianqian Liu, Ruofeng Tang, Xingyue Huang, Bing Zhao, Mengyang Zhang, Youqiang Zhou

专题命中 视觉推理 :vision-language model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04088 2025-08-08 cs.CL 50%

GM-PRM: A Generative Multimodal Process Reward Model for Multimodal Mathematical Reasoning

Jianghangfan Zhang, Yibo Yan, Kening Zheng, Xin Zou, Song Dai, Xuming Hu

专题命中 视觉推理 :multimodal large language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视觉定位与Grounding 8 篇

2508.05021 2025-08-08 cs.RO 82%

MAG-Nav: Language-Driven Object Navigation Leveraging Memory-Reserved Active Grounding

Weifan Zhang, Tingguang Li, Yuzhen Liu

机构 * Tencent Robotics X(腾讯机器人X) Harbin Institute of Technology(哈尔滨工业大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);visual language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05323 2025-08-08 cs.CV 70%

Textual Inversion for Efficient Adaptation of Open-Vocabulary Object Detectors Without Forgetting

Frank Ruis, Gertjan Burghouts, Hugo Kuijf

机构 * TNO(荷兰技术院) Intelligent Imaging(智能成像)

专题命中 视觉定位与Grounding :vision language model(abstract);VLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05409 2025-08-08 cs.CV cs.SD eess.AS 57%

From Detection to Correction: Backdoor-Resilient Face Recognition via Vision-Language Trigger Detection and Noise-Based Neutralization

Farah Wahida, M. A. P. Chamikara, Yashothara Shanmugarasa, Mohan Baruwal Chhetri, Thilina Ranbaduge, Ibrahim Khalil

机构 * RMIT University, Australia(皇家墨尔本理工大学)

专题命中 视觉定位与Grounding :vision language model(abstract);分类 cs.CV

Comments 19 Pages, 24 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04868 2025-08-08 cs.CV 57%

Dual-Stream Attention with Multi-Modal Queries for Object Detection in Transportation Applications

Noreen Anwar, Guillaume-Alexandre Bilodeau, Wassim Bouachir

机构 * LITIV, Polytechnique Montréal(Polytechnique Montréal 的 LITIV) Data Science Laboratory, Université du Québec (TELUQ)(Université du Québec (TELUQ) 的 Data Science Laboratory)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05534 2025-08-08 cs.CL 50%

CoCoLex: Confidence-guided Copy-based Decoding for Grounded Legal Text Generation

Santosh T. Y. S. S, Youssef Tarek Elkhayat, Oana Ichim, Pranav Shetty, Dongsheng Wang, Zhiqiang Ma, Armineh Nourbakhsh, Xiaomo Liu

机构 * School of Computation, Information, and Technology, Technical University of Munich(计算、信息与技术学院,慕尼黑技术大学) Graduate Institute of International and Development Studies, Geneva(国际与发展研究研究生院,日内瓦) JPMorgan AI Research(摩根大通人工智能研究)

专题命中 视觉定位与Grounding :grounding(abstract)

Comments Accepted to ACL 2025-Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05061 2025-08-08 cs.DB cs.IR 50%

Data-Aware Socratic Query Refinement in Database Systems

Ruiyuan Zhang, Chrysanthi Kosyfaki, Xiaofang Zhou

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02068 2025-08-08 cs.RO 50%

"Set It Up": Functional Object Arrangement with Compositional Generative Models (Journal Version)

Yiqing Xu, Jiayuan Mao, Linfeng Li, Yilun Du, Tomas Lozáno-Pérez, Leslie Pack Kaelbling, David Hsu

机构 * School of Computing, National University of Singapore(新加坡国立大学计算机学院) CSAIL, Massachusetts Institute of Technology(麻省理工学院计算机科学与人工智能实验室)

专题命中 视觉定位与Grounding :grounding(abstract)

Comments This is the journal version accepted to the International Journal of Robotics Research (IJRR). It extends our prior work presented at Robotics: Science and Systems (RSS) 2024, with a new compositional program induction pipeline from natural language, and expanded evaluations on personalized bookshelf and bedroom furniture layout tasks

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05236 2025-08-08 cs.MA 50%

Towards Language-Augmented Multi-Agent Deep Reinforcement Learning

Maxime Toquebiau, Jae-Yun Jun, Faïz Benamar, Nicolas Bredeche

专题命中 视觉定位与Grounding :grounding(abstract)

Comments Accespted at the European Conference on Artificial Intelligence 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 幻觉与鲁棒性 6 篇

2508.05167 2025-08-08 cs.CV 83%

PhysPatch: A Physically Realizable and Transferable Adversarial Patch Attack for Multimodal Large Language Models-based Autonomous Driving Systems

Qi Guo, Xiaojun Jia, Shanmin Pang, Simeng Qin, Lin Wang, Ju Jia, Yang Liu, Qing Guo

专题命中 幻觉与鲁棒性 :multimodal large language model(title,abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04942 2025-08-08 cs.CV 83%

Accelerating Conditional Prompt Learning via Masked Image Modeling for Vision-Language Models

Phuoc-Nguyen Bui, Khanh-Binh Nguyen, Hyunseung Choo

机构 * Convergence Research Institute, Sungkyunkwan University(convergence research institute, 首尔大学) School of Information Technology, Deakin University(信息科技学院, 德金大学) Department of Electrical and Computer Engineering, Sungkyunkwan University(电气与计算机工程系, 首尔大学)

专题命中 幻觉与鲁棒性 :vision-language model(title,abstract);VLM(abstract);分类 cs.CV

Comments ACMMM-LAVA 2025, 10 pages, camera-ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05237 2025-08-08 cs.CV cs.AI 81%

Navigating the Trade-off: A Synthesis of Defensive Strategies for Zero-Shot Adversarial Robustness in Vision-Language Models

Zane Xu, Jason Sun

专题命中 幻觉与鲁棒性 :vision-language model(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05087 2025-08-08 cs.MM cs.AI cs.CL cs.CR 79%

JPS: Jailbreak Multimodal Large Language Models with Collaborative Visual Perturbation and Textual Steering

Renmiao Chen, Shiyao Cui, Xuancheng Huang, Chengwei Pan, Victor Shea-Jay Huang, QingLin Zhang, Xuan Ouyang, Zhexin Zhang, Hongning Wang, Minlie Huang

机构 * CoAI group, DCST, Tsinghua University(清华大学DCST学院) Beihang University(北航大学)

专题命中 幻觉与鲁棒性 :multimodal large language model(title,abstract);分类 cs.AI

Comments 10 pages, 3 tables, 2 figures, to appear in the Proceedings of the 33rd ACM International Conference on Multimedia (MM '25)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05083 2025-08-08 cs.AI 79%

MedMKEB: A Comprehensive Knowledge Editing Benchmark for Medical Multimodal Large Language Models

Dexuan Xu, Jieyi Wang, Zhongyan Chai, Yongzhi Cao, Hanpin Wang, Huamin Zhang, Yu Huang

专题命中 幻觉与鲁棒性 :multimodal large language model(title,abstract);分类 cs.AI

Comments 18 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05527 2025-08-08 cs.CV 57%

AI vs. Human Moderators: A Comparative Evaluation of Multimodal LLMs in Content Moderation for Brand Safety

Adi Levi, Or Levi, Sardhendu Mishra, Jonathan Morra

机构 * Zefr Inc(Zefr公司)

专题命中 幻觉与鲁棒性 :multimodal large language model(abstract);分类 cs.CV

Comments Accepted to the Computer Vision in Advertising and Marketing (CVAM) workshop at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

5. VLM训练与架构 7 篇

2508.05602 2025-08-08 cs.CV 89%

LLaVA-RE: Binary Image-Text Relevancy Evaluation with Multimodal Large Language Model

Tao Sun, Oliver Liu, JinJin Li, Lan Ma

机构 * Stony Brook University(斯通布罗克大学) Amazon(亚马逊)

专题命中 VLM训练与架构 :LLaVA(title,abstract);multimodal large language model(title,abstract);MLLM(abstract);分类 cs.CV

Comments Published in the First Workshop of Evaluation of Multi-Modal Generation 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05547 2025-08-08 cs.LG cs.AI cs.CV 85%

Adapting Vision-Language Models Without Labels: A Comprehensive Survey

Hao Dong, Lijun Sheng, Jian Liang, Ran He, Eleni Chatzi, Olga Fink

机构 * ETH Zürich(苏黎世联邦理工学院) University of Science and Technology of China(中国科学技术大学) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) EPFL(苏黎世联邦理工学院)

专题命中 VLM训练与架构 :vision-language model(title,abstract);VLM(abstract);分类 cs.CV、cs.AI、cs.LG

Comments Discussions, comments, and questions are welcome in \url{https://github.com/tim-learn/Awesome-LabelFree-VLMs}

详情

展开后加载摘要…

URL PDF HTML 收藏