arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-09-16 至 2025-09-16 共收录 38 信号源:cs.CV, cs.AI, cs.LG

1. 视觉问答 6 篇

2509.11862 2025-09-16 cs.CV cs.AI cs.LG 91%

Bridging Vision Language Models and Symbolic Grounding for Video Question Answering

Haodi Ma, Vyom Pathak, Daisy Zhe Wang

机构 * Univerisy of Florida(佛罗里达大学)

专题命中 视觉问答 :vision language model(title,abstract);grounding(title,abstract);VLM(abstract);InternVL(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11986 2025-09-16 cs.CV cs.CL 74%

Lost in Embeddings: Information Loss in Vision-Language Models

Wenyan Li, Raphael Tang, Chengzu Li, Caiqi Zhang, Ivan Vulić, Anders Søgaard

机构 * University of Copenhagen(哥本哈根大学) Microsoft(微软) University of Cambridge(剑桥大学)

专题命中 视觉问答 :vision-language model(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10129 2025-09-16 cs.CL cs.IR 67%

Towards Reliable and Interpretable Document Question Answering via VLMs

Alessio Chen, Simone Giovannini, Andrea Gemelli, Fabio Coppini, Simone Marinai

机构 * Università degli Studi di Firenze(佛罗伦萨大学)

专题命中 视觉问答 :vision-language model(abstract);grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.07030 2025-09-16 cs.CL cs.AI cs.CV cs.IR cs.LG 67%

FM2DS: Few-Shot Multimodal Multihop Data Synthesis with Knowledge Distillation for Question Answering

Amirhossein Abaskohi, Spandana Gella, Giuseppe Carenini, Issam H. Laradji

机构 * Department of Computer Science(计算机科学系) The University of British Columbia(不列颠哥伦比亚大学) ServiceNow Research(ServiceNow研究)

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.AI、cs.LG

Comments Findings of EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11589 2025-09-16 cs.CV 57%

MVQA-68K: A Multi-dimensional and Causally-annotated Dataset with Quality Interpretability for Video Assessment

Yanyun Pu, Kehan Li, Zeyi Huang, Zhijie Zhong, Kaixiang Yang

机构 * Huawei Technologies Co.(华为技术有限公司) South China University of Technology(南方科技大学)

专题命中 视觉问答 :multimodal large language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.05479 2025-09-16 cs.CV 57%

LATTE: Learning to Think with Vision Specialists

Zixian Ma, Jianguo Zhang, Zhiwei Liu, Jieyu Zhang, Juntao Tan, Manli Shu, Juan Carlos Niebles, Shelby Heinecke, Huan Wang, Caiming Xiong, Ranjay Krishna, Silvio Savarese

机构 * University of Washington(华盛顿大学) Salesforce Research(Salesforce研究)

专题命中 视觉问答 :vision-language model(abstract);分类 cs.CV

Journal ref EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 视觉推理 4 篇

2508.13229 2025-09-16 cs.LG cs.CV 84%

RISE: Enhancing VLM Image Annotation with Self-Supervised Reasoning

Suhang Hu, Wei Hu, Yuhang Su, Fan Zhang

专题命中 视觉推理 :VLM(title,abstract);vision-language model(abstract);分类 cs.CV、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12132 2025-09-16 cs.CV cs.CL 83%

Look Again, Think Slowly: Enhancing Visual Reflection in Vision-Language Models

Pu Jian, Junhong Wu, Wei Sun, Chen Wang, Shuo Ren, Jiajun Zhang

专题命中 视觉推理 :vision-language model(title,abstract);visual reasoning(abstract);分类 cs.CV

Comments EMNLP2025 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.09070 2025-09-16 cs.LG cs.AI cs.CV 78%

FairCoT: Enhancing Fairness in Text-to-Image Generation via Chain of Thought Reasoning with Multimodal Large Language Models

Zahraa Al Sahili, Ioannis Patras, Matthew Purver

机构 * School of Electronic Engineering and Computer Science, Queen Mary University of London(伦敦女王学院电子工程与计算机科学学院) Department of Knowledge Technologies, Jožef Stefan Institute(Jožef Stefan研究所知识技术系)

专题命中 视觉推理 :multimodal large language model(title);分类 cs.CV、cs.AI、cs.LG

Comments Accepted at EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00284 2025-09-16 cs.RO cs.AI 70%

LightEMMA: Lightweight End-to-End Multimodal Model for Autonomous Driving

Zhijie Qiao, Haowei Li, Zhong Cao, Henry X. Liu

机构 * Department of Civil and Environmental Engineering, University of Michigan(土木与环境工程系,密歇根大学) University of Michigan Transportation Research Institute(密歇根大学交通研究所)

专题命中 视觉推理 :vision-language model(abstract);VLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视觉定位与Grounding 9 篇

2509.10345 2025-09-16 cs.CV cs.AI 88%

Towards Understanding Visual Grounding in Visual Language Models

Georgios Pantazopoulos, Eda B. Özyiğit

机构 * The Alan Turing Institute(艾伦·图灵研究所) Heriot-Watt University(赫瑞-沃德大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);visual language model(title);vision language model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.02573 2025-09-16 cs.CV 83%

Remote Sensing SpatioTemporal Vision-Language Models: A Comprehensive Survey

Chenyang Liu, Jiafan Zhang, Keyan Chen, Man Wang, Zhengxia Zou, Zhenwei Shi

机构 * Department of Aerospace Intelligent Science and Technology, School of Astronautics, Beihang University(航空航天智能科学与技术系,航天学院,北航) Key Laboratory of Spacecraft Design Optimization and Dynamic Simulation Technologies, Ministry of Education, China(航天器设计优化与动态仿真技术重点实验室,教育部,中国) College of Computer Science, Inner Mongolia University(计算机科学学院,内蒙古大学) Shen Yuan Honors College of Beihang University(盛元荣誉学院,北航)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);grounding(abstract);分类 cs.CV

Comments Published in IEEE Geoscience and Remote Sensing Magazine

Journal ref IEEE Geoscience and Remote Sensing Magazine, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11866 2025-09-16 cs.CV 79%

Dr.V: A Hierarchical Perception-Temporal-Cognition Framework to Diagnose Video Hallucination by Fine-grained Spatial-Temporal Grounding

Meng Luo, Shengqiong Wu, Liqiang Jing, Tianjie Ju, Li Zheng, Jinxiang Lai, Tianlong Wu, Xinya Du, Jian Li, Siyuan Yan, Jiebo Luo, William Yang Wang, Hao Fei, Mong-Li Lee, Wynne Hsu

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments 25 pages, 16 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.10424 2025-09-16 cs.CV cs.AI 79%

What is the Visual Cognition Gap between Humans and Multimodal LLMs?

Xu Cao, Yifan Shen, Bolin Lai, Wenqian Ye, Yunsheng Ma, Joerg Heintz, Jintai Chen, Meihuan Huang, Jianguo Cao, Aidong Zhang, James M. Rehg

机构 * Department of Computer Science, University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校计算机科学系) College of Computing, Georgia Institute of Technology(佐治亚理工学院计算机学院) Department of Computer Science, University of Virginia(弗吉尼亚大学计算机科学系) Digital Twin Lab, Purdue University(普渡大学数字孪生实验室) HKUST (Guangzhou)(香港科技大学(广州)) Department of Rehabilitation Medicine, Shenzhen Children’s Hospital(深圳儿童医院康复医学系)

专题命中 视觉定位与Grounding :vision language model(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments COLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12145 2025-09-16 cs.CV 74%

Open-ended Hierarchical Streaming Video Understanding with Vision Language Models

Hyolim Kang, Yunsu Park, Youngbeom Yoo, Yeeun Choi, Seon Joo Kim

机构 * Yonsei University(延世大学)

专题命中 视觉定位与Grounding :vision language model(title);分类 cs.CV

Comments 17 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11840 2025-09-16 cs.CV 57%

Synthetic Captions for Open-Vocabulary Zero-Shot Segmentation

Tim Lebailly, Vijay Veerabadran, Satwik Kottur, Karl Ridgeway, Michael Louis Iuzzolino

机构 * Meta KU Leuven(鲁汶大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments ICCV 2025 CDEL Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11714 2025-09-16 eess.IV cs.LG 57%

EMeRALDS: Electronic Medical Record Driven Automated Lung Nodule Detection and Classification in Thoracic CT Images

Hafza Eman, Furqan Shaukat, Muhammad Hamza Zafar, Syed Muhammad Anwar

机构 * Faculty of Electrical and Electronics Engineering, University of Engineering(电气电子工程学院,工程大学) Department of Engineering Sciences, University of Agder(工程科学系,阿格德大学) Sheikh Zayed Institute for Pediatric Surgical Innovation, Children’s National Hospital(谢赫扎耶德小儿外科创新研究所,儿童医院) School of Medicine and Health Sciences, George Washington University(医学与健康科学学院,乔治华盛顿大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11336 2025-09-16 cs.AI 57%

The power of dynamic causality in observer-based design for soft sensor applications

William Farlessyost, Sebastian Oberst, Shweta Singh

机构 * organization= Agricultural \& Biological Engineering, Purdue University , country= USA organization= Environmental \& Ecological Engineering, Purdue University , country= USA organization= Davidson School of Chemical Engineering, Purdue University , country= USA organization= Mechanical \& Mechatronic Engineering, University of Technology Sydney (UTS) , country= Australia

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10478 2025-09-16 cs.NI cs.LG cs.SY eess.SY 57%

The LLM as a Network Operator: A Vision for Generative AI in the 6G Radio Access Network

Oluwaseyi Giwa, Michael Adewole, Tobi Awodumila, Pelumi Aderinto

机构 * African Institute for Mathematical Sciences(非洲数学科学研究所)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

Comments Submitted to Workshop on AI and ML for Next-Generation Wireless Communications and Networking, NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

4. GUI与屏幕智能体 3 篇

2503.01062 2025-09-16 cs.LG cs.AI 84%

Offline RLAIF: Piloting VLM Feedback for RL via SFO

Jacob Beck

机构 * Jacob Beck 1

专题命中 GUI与屏幕智能体 :VLM(title,abstract);vision-language model(abstract);分类 cs.AI、cs.LG

Comments Code is provided at https://github.com/jacooba/OfflineRLAIF

Journal ref Published at The RLC 2025 Workshop on Reinforcement Learning Beyond Rewards: Ingredients for Developing Generalist Agents

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.17406 2025-09-16 cs.CV 70%

Seeing the Undefined: Chain-of-Action for Generative Semantic Labels

Meng Wei, Zhongnian Li, Peng Ying, Xinzheng Xu

机构 * 1 School of Computer Science Technology / School of Artificial Intelligence, China University of Mining Technology Xuzhou China 2 Mine Digitization Engineering Research Center of the Ministry of Education Xuzhou China Technology Xuzhou China 2 The State Key Laboratory of CAD\&CG, Zhejiang University Hangzhou China 3 Mine Digitization Engineering Research Center of the Ministry of Education Xuzhou China 2 Mine Digitization Engineering Research Center of the Ministry of Education 2 The State Key Laboratory of CAD\&CG, Zhejiang University 3 Mine Digitization Engineering Research Center of the Ministry of Education

专题命中 GUI与屏幕智能体 :vision-language model(abstract);VLM(abstract);分类 cs.CV

Comments 15 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10884 2025-09-16 cs.RO cs.CV 57%

Nav-R1: Reasoning and Navigation in Embodied Scenes

Qingxiang Liu, Ting Huang, Zeyu Zhang, Hao Tang

机构 * Shanghai University of Engineering Science(上海工程技术大学) Peking University(北京大学)

专题命中 GUI与屏幕智能体 :grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 幻觉与鲁棒性 6 篇

2408.14435 2025-09-16 cs.CV cs.AI cs.CY cs.LG 82%

Social Perception of Faces in a Vision-Language Model

Carina I. Hausladen, Manuel Knott, Colin F. Camerer, Pietro Perona

机构 * California Institute of Technology(加州理工学院) ETH Zurich(苏黎世联邦理工学院) Computational Social Science(计算社会科学) Swiss Data Science Center(瑞士数据科学中心) Empa(瑞士联邦材料科学与技术研究院)

专题命中 幻觉与鲁棒性 :vision-language model(title,abstract);分类 cs.CV、cs.AI、cs.LG

Journal ref Published in the Proceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency (FAccT 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11287 2025-09-16 cs.CV cs.CL 79%

Mitigating Hallucinations in Large Vision-Language Models by Self-Injecting Hallucinations

Yifan Lu, Ziqi Zhang, Chunfeng Yuan, Jun Gao, Congxuan Zhang, Xiaojuan Qi, Bing Li, Weiming Hu

机构 * Beijing Key Laboratory of Super Intelligent Security of Multi-Modal Information, CASIA(北京多模态信息超级智能安全重点实验室,中国科学院自动化所) State Key Laboratory of Multimodal Artificial Intelligence Systems, CASIA(多模态人工智能系统国家重点实验室,中国科学院自动化所) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Hello Group(Hello集团) Nanchang Hangkong University(南昌航空大学) The University of Hong Kong(香港大学) School of Information Science and Technology, ShanghaiTech University(上海科技大学信息科学与技术学院)

专题命中 幻觉与鲁棒性 :vision-language model(title,abstract);分类 cs.CV

Comments emnlp 2025 accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16146 2025-09-16 cs.CV cs.AI cs.CL cs.LG 67%

Steering LVLMs via Sparse Autoencoder for Hallucination Mitigation

Zhenglin Hua, Jinghan He, Zijun Yao, Tianxu Han, Haiyun Guo, Yuheng Jia, Junfeng Fang

机构 * School of Computer Science and Engineering, Southeast University(东南大学计算机科学与工程学院) Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications (Southeast University)(东南大学新一代人工智能技术及其交叉应用关键实验室) Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所基础模型研究中心) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系) Wuhan University of Technology(武汉理工大学) National University of Singapore(新加坡国立大学)

专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.CV、cs.AI、cs.LG

Comments Accepted to Findings of EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06336 2025-09-16 cs.CV cs.AI cs.CR 62%

Multi-View Slot Attention Using Paraphrased Texts for Face Anti-Spoofing

Jeongmin Yu, Susang Kim, Kisu Lee, Taekyoung Kwon, Won-Yong Shin, Ha Young Kim

机构 * Yonsei University(延世大学)

专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.CV、cs.AI

Comments Accepted to ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01064 2025-09-16 cs.CV cs.AI 62%

Fighting Fire with Fire (F3): A Training-free and Efficient Visual Adversarial Example Purification Method in LVLMs

Yudong Zhang, Ruobing Xie, Yiqing Huang, Jiansheng Chen, Xingwu Sun, Zhanhui Kang, Di Wang, Yu Wang

机构 * Tsinghua University, Tencent(清华大学,腾讯) Tencent(腾讯) University of Science and Technology Beijing(北京科技大学) Tencent, University of Macau(腾讯,澳门大学) Tsinghua University(清华大学)

专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.CV、cs.AI

Comments Accepted by ACM Multimedia 2025 BNI track (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10570 2025-09-16 cs.RO cs.AI 57%

Large Foundation Models for Trajectory Prediction in Autonomous Driving: A Comprehensive Survey

Wei Dai, Shengen Wu, Wei Wu, Zhenhao Wang, Sisuo Lyu, Haicheng Liao, Limin Yu, Weiping Ding, Runwei Guan, Yutao Yue

机构 * Department of Mathematical Sciences, School of Physical sciences, University of Liverpool(利物浦大学数学科学系) Department of Communications and Networking, School of Advanced Technology, Xi’an Jiaotong-Liverpool University(西安交通大学利物浦大学通讯与网络系) Thrust of Artificial Intelligence, The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)人工智能方向) Deep Interdisciplinary Intelligence Lab, The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)深度跨学科智能实验室) School of Mathematics and Statistics, Shandong University(山东大学数学与统计学院) School of Artificial Intelligence and Computer Science, Nantong University(南通大学人工智能与计算机科学学院) Thrust of Data Science and Analytics, The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)数据科学与分析方向) Institute of Deep Perception Technology, Jiangsu(江苏深度感知技术研究院)

专题命中 幻觉与鲁棒性 :multimodal large language model(abstract);分类 cs.AI

Comments 22 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

6. VLM训练与架构 6 篇

2509.11961 2025-09-16 cs.CL 89%

Spec-LLaVA: Accelerating Vision-Language Models with Dynamic Tree-Based Speculative Decoding

Mingxiao Huo, Jiayi Zhang, Hewei Wang, Jinfeng Xu, Zheyu Chen, Huilin Tai, Yijun Chen

机构 * Carnegie Mellon University(卡内基梅隆大学) University of Nottingham(诺丁汉大学) The University of Hong Kong(香港大学) The Hong Kong Polytechnic University(香港理工大学) Columbia University(哥伦比亚大学) University of California, Berkeley(加州大学伯克利分校)

专题命中 VLM训练与架构 :vision-language model(title,abstract);LLaVA(title,abstract);VLM(abstract)

Comments 7pages, accepted by ICML TTODLer-FM workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.15688 2025-09-16 cs.CL cs.AI cs.LG 73%

Transformer-Based Multimodal Knowledge Graph Completion with Link-Aware Contexts

Haodi Ma, Dzmitry Kasinets, Daisy Zhe Wang

机构 * Department of Computer and Information Science and Engineering, University of Florida(计算机与信息科学与工程系,佛罗里达大学)

专题命中 VLM训练与架构 :vision-language model(abstract);VLM(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏