arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-11-14 至 2025-11-14 共收录 38 信号源:cs.CV, cs.AI, cs.LG

1. 视觉问答 5 篇

2511.10059 2025-11-14 cs.CV 81%

When Eyes and Ears Disagree: Can MLLMs Discern Audio-Visual Confusion?

Qilang Ye, Wei Zeng, Meng Liu, Jie Zhang, Yupeng Hu, Zitong Yu, Yu Zhou

专题命中 视觉问答 :visual reasoning(abstract);visual question answering(abstract);multimodal large language model(abstract);MLLM(abstract)

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10591 2025-11-14 cs.CL cs.AI 74%

Mined Prompting and Metadata-Guided Generation for Wound Care Visual Question Answering

Bavana Durgapraveen, Sornaraj Sivasankaran, Abhinand Balachandran, Sriram Rajkumar

机构 * EXL Health AI Lab at MEDIQA-WV 2025(EXL健康AI实验室)

专题命中 视觉问答 :visual question answering(title);分类 cs.AI

Comments 2 figures, 11 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09651 2025-11-14 cs.IR cs.AI cs.CL cs.LG eess.SP 62%

Retrieval-Augmented Generation for Reliable Interpretation of Radio Regulations

Zakaria El Kassimi, Fares Fourati, Mohamed-Slim Alouini

机构 * KAUST(卡斯泰尔大学)

专题命中 视觉问答 :grounding(abstract);分类 cs.AI、cs.LG

Comments 12 pages, 7 figures, AI4NextG @ NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09868 2025-11-14 cs.CV 57%

Remember Me: Bridging the Long-Range Gap in LVLMs with Three-Step Inference-Only Decay Resilience Strategies

Peng Gao, Yujian Lee, Xiaofeng Zhang, Zailong Chen, Hui Zhang

专题命中 视觉问答 :vision-language model(abstract);分类 cs.CV

Comments Accepted in AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04369 2025-11-14 cs.CV 57%

TSPO: Temporal Sampling Policy Optimization for Long-form Video Language Understanding

Canhui Tang, Zifan Han, Hongbo Sun, Sanping Zhou, Xuchong Zhang, Xin Wei, Ye Yuan, Huayu Zhang, Jinglin Xu, Hao Sun

专题命中 视觉问答 :multimodal large language model(abstract);分类 cs.CV

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 视觉推理 6 篇

2511.10279 2025-11-14 cs.CV 85%

PROPA: Toward Process-level Optimization in Visual Reasoning via Reinforcement Learning

Yanbei Jiang, Chao Lei, Yihao Ding, Krista Ehinger, Jey Han Lau

机构 * University of Melbourne(墨尔本大学)

专题命中 视觉推理 :visual reasoning(title,abstract);vision-language model(abstract);VLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10017 2025-11-14 cs.CV 85%

AffordBot: 3D Fine-grained Embodied Reasoning via Multimodal Large Language Models

Xinyi Wang, Xun Yang, Yanlong Xu, Yuchen Wu, Zhen Li, Na Zhao

机构 * University of Science and Technology of China(科学技术大学) Singapore University of Technology and Design(新加坡科技设计大学) Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))

专题命中 视觉推理 :multimodal large language model(title,abstract);grounding(abstract);MLLM(abstract);分类 cs.CV

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21572 2025-11-14 cs.CL 82%

Aligning MLLM Benchmark With Human Preferences via Structural Equation Modeling

Shengwu. Xiong, Tianyu. Zou, Cong. Wang, Xuelong Li

机构 * Interdisciplinary Artificial Intelligence Research Institute, Wuhan College(交叉学科人工智能研究 institute,武汉学院) School of Computer and Artificial Intelligence, Wuhan University of Technology(计算机与人工智能学院,武汉理工大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Sanya Science and Education Innovation Park, Wuhan University of Technology(三亚科学教育创新园,武汉理工大学) Institute of Automation, Chinese Academy of Sciences(自动化研究所,中国科学院) School of Mathematics and Statistics, Northwestern Polytechnical University(数学与统计学院,西北工业大学) Institute of Artificial Intelligence (TeleAI) of China Telecom(中国电信人工智能研究所(TeleAI))

专题命中 视觉推理 :MLLM(title,abstract);multimodal large language model(abstract)

Comments 12 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21538 2025-11-14 cs.CV cs.AI 79%

Caption This, Reason That: VLMs Caught in the Middle

Zihan Weng, Lucas Gomez, Taylor Whittington Webb, Pouya Bashivan

机构 * Integrated Program in Neuroscience (IPN) McGill University(神经科学联合计划 麦吉尔大学) Mila, University of Montreal(蒙特利尔大学Mila) Microsoft Research USA(微软研究院美国总部) Department of Physiology McGill University(生理学系 麦吉尔大学)

专题命中 视觉推理 :vision-language model(abstract);VLM(abstract);visual reasoning(abstract);分类 cs.CV、cs.AI

Comments Paper accepted by nips 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07250 2025-11-14 cs.CV cs.AI 62%

MVU-Eval: Towards Multi-Video Understanding Evaluation for Multimodal LLMs

Tianhao Peng, Haochen Wang, Yuanxing Zhang, Zekun Wang, Zili Wang, Gavin Chang, Jian Yang, Shihao Li, Yanghai Wang, Xintao Wang, Houyi Li, Wei Ji, Pengfei Wan, Steven Huang, Zhaoxiang Zhang, Jiaheng Liu

机构 * Nanjing University(南京大学) CASIA(中国科学院自动化研究所) Kuaishou Technology(快手科技) M-A-P

专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.CV、cs.AI

Journal ref The Thirty-Ninth Annual Conference on Neural Information Processing Systems (NeurIPS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10648 2025-11-14 cs.CV 57%

Enhancing the Outcome Reward-based RL Training of MLLMs with Self-Consistency Sampling

Jiahao Wang, Weiye Xu, Aijun Yang, Wengang Zhou, Lewei Lu, Houqiang Li, Xiaohua Wang, Jinguo Zhu

机构 * Xi’an Jiaotong University(西安交通大学) University of Science and Technology of China(中国科学技术大学) SenseTime Research(商汤科技研究院)

专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.CV

Comments Accepted to NeurIPS 2025 (The Thirty-Ninth Annual Conference on Neural Information Processing Systems)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视觉定位与Grounding 7 篇

2507.19110 2025-11-14 cs.CV 83%

LISA: A Layer-wise Integration and Suppression Approach for Hallucination Mitigation in Multimodal Large Language Models

Zhihui Guo, Xin Man, Hui Xu, Jie Shao, Zhiguo Jiang, Xianchao Zhang, Heng Tao Shen

专题命中 视觉定位与Grounding :multimodal large language model(title,abstract);grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05615 2025-11-14 cs.CV cs.AI cs.CL 81%

Test-Time Reinforcement Learning for GUI Grounding via Region Consistency

Yong Du, Yuchen Yan, Fei Tang, Zhengxi Lu, Chang Zong, Weiming Lu, Shengpei Jiang, Yongliang Shen

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.AI

Comments [Accepted by AAAI2026] Project Page: https://zju-real.github.io/gui-rcpo Code: https://github.com/zju-real/gui-rcpo

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09958 2025-11-14 cs.CV cs.AI 73%

Zero-Shot Referring Expression Comprehension via Vison-Language True/False Verification

Jeffrey Liu, Rongbin Hu

专题命中 视觉定位与Grounding :VLM(abstract);grounding(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09955 2025-11-14 cs.CV 70%

Robust Object Detection with Pseudo Labels from VLMs using Per-Object Co-teaching

Uday Bhaskar, Rishabh Bhattacharya, Avinash Patel, Sarthak Khoche, Praveen Anil Kulkarni, Naresh Manwani

机构 * Machine Learning Lab IIIT Hyderabad(IIIT Hyderabad 机器学习实验室) Bosch Global Software Technologies(博世全球软件技术公司)

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10020 2025-11-14 cs.CV cs.AI 62%

Anomagic: Crossmodal Prompt-driven Zero-shot Anomaly Generation

Yuxin Jiang, Wei Luo, Hui Zhang, Qiyu Chen, Haiming Yao, Weiming Shen, Yunkang Cao

专题命中 视觉定位与Grounding :multimodal large language model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10518 2025-11-14 cs.CV cs.RO 57%

SemanticVLA: Semantic-Aligned Sparsification and Enhancement for Efficient Robotic Manipulation

Wei Li, Renshan Zhang, Rui Shao, Zhijian Fang, Kaiwen Zhou, Zhuotao Tian, Liqiang Nie

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments Accepted to AAAI 2026 (Oral), Project Page: https://github.com/JiuTian-VL/SemanticVLA

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23740 2025-11-14 cs.CV cs.GR 57%

LayerPeeler: Autoregressive Peeling for Layer-wise Image Vectorization

Ronghuan Wu, Wanchao Su, Jing Liao

机构 * City University of Hong Kong(香港城市大学) Monash University(墨尔本大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments Project Page: https://layerpeeler.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 文档图表理解 2 篇

2511.10552 2025-11-14 cs.CL 67%

URaG: Unified Retrieval and Generation in Multimodal LLMs for Efficient Long Document Understanding

Yongxin Shi, Jiapeng Wang, Zeyu Shan, Dezhi Peng, Zening Lin, Lianwen Jin

专题命中 文档图表理解 :multimodal large language model(abstract);MLLM(abstract)

Comments Accepted by AAAI 2026 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09919 2025-11-14 cs.CV 57%

MosaicDoc: A Large-Scale Bilingual Benchmark for Visually Rich Document Understanding

Ketong Chen, Yuhao Chen, Yang Xue

机构 * Ketong Chen, Yuhao Chen, Yang Xue(作者)

专题命中 文档图表理解 :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

5. GUI与屏幕智能体 2 篇

2508.05294 2025-11-14 cs.RO cs.AI cs.LG 81%

Towards Embodied Agentic AI: Review and Classification of LLM- and VLM-Driven Robot Autonomy and Interaction

Sahar Salimpour, Lei Fu, Kajetan Rachwał, Pascal Bertrand, Kevin O'Sullivan, Robert Jakob, Farhad Keramat, Leonardo Militano, Giovanni Toffetti, Harry Edelman, Jorge Peña Queralta

机构 * Department of Computing, University of Turku(图尔库大学计算机系) Institute of Computer Science, Zurich University of Applied Sciences(应用科学大学计算机科学研究所) Centre for Artificial Ingelligence, Zurich University of Applied Sciences(应用科学大学人工智能中心) Agentic Systems Lab, Department of Management, Technology and Economics, ETH Zürich(苏黎世联邦理工学院管理、科技与经济系代理系统实验室) Faculty of Mathematics and Information Science, Warsaw University of Technology(华沙技术大学数学与信息科学学院)

专题命中 GUI与屏幕智能体 :VLM(title);vision-language model(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10615 2025-11-14 cs.CV cs.CL 57%

Towards Blind and Low-Vision Accessibility of Lightweight VLMs and Custom LLM-Evals

Shruti Singh Baghel, Yash Pratap Singh Rathore, Sushovan Jena, Anurag Pradhan, Amit Shukla, Arnav Bhavsar, Pawan Goyal

机构 * Indian Institute of Technology Mandi(印度理工学院曼迪分校) Vellore Institute of Technology(韦洛雷理工学院) Indian Institute of Technology Kharagpur(印度理工学院哈里科普分校)

专题命中 GUI与屏幕智能体 :vision-language model(abstract);分类 cs.CV

Comments 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 幻觉与鲁棒性 5 篇

2511.10074 2025-11-14 cs.CV cs.SY eess.SY 70%

VLF-MSC: Vision-Language Feature-Based Multimodal Semantic Communication System

Gwangyeon Ahn, Jiwan Seo, Joonhyuk Kang

专题命中 幻觉与鲁棒性 :vision-language model(abstract);VLM(abstract);分类 cs.CV

Comments To appear in the AI4NextG Workshop at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10268 2025-11-14 cs.AI 57%

Causal-HalBench: Uncovering LVLMs Object Hallucinations Through Causal Intervention

Zhe Xu, Zhicai Wang, Junkang Wu, Jinda Lu, Xiang Wang

专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.AI

Comments accepted for publication in the Association for the Advancement of Artificial Intelligence (AAAI), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15721 2025-11-14 cs.CL cs.AI 57%

EcomMMMU: Strategic Utilization of Visuals for Robust Multimodal E-commerce Models

Xinyi Ling, Hanwen Du, Zhihui Zhu, Xia Ning

机构 * Department of Computer Science and Engineering, The Ohio State University(俄亥俄州立大学计算机科学与工程系) Translational Data Analytics Institute, The Ohio State University(俄亥俄州立大学转化数据分析研究所) Department of Biomedical Informatics, The Ohio State University(俄亥俄州立大学生物医学信息学系)

专题命中 幻觉与鲁棒性 :multimodal large language model(abstract);分类 cs.AI

Comments ICJNLP-AACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.18638 2025-11-14 cs.CR cs.AI cs.CL 57%

Graph of Attacks with Pruning: Optimizing Stealthy Jailbreak Prompt Generation for Enhanced LLM Content Moderation

Daniel Schwartz, Dmitriy Bespalov, Zhe Wang, Ninad Kulkarni, Yanjun Qi

机构 * Amazon Bedrock Science(亚马逊Bedrock科学) Drexel University(德雷塞尔大学) University of Virginia(弗吉尼亚大学)

专题命中 幻觉与鲁棒性 :VLM(abstract);分类 cs.AI

Comments 14 pages, 5 figures; published in EMNLP 2025 ; Code at: https://github.com/dsbuddy/GAP-LLM-Safety

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10075 2025-11-14 cs.CL 50%

Format Matters: The Robustness of Multimodal LLMs in Reviewing Evidence from Tables and Charts

Xanh Ho, Yun-Ang Wu, Sunisth Kumar, Florian Boudin, Atsuhiro Takasu, Akiko Aizawa

专题命中 幻觉与鲁棒性 :multimodal large language model(abstract)

Comments Accepted at AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏

7. VLM训练与架构 9 篇

2511.09973 2025-11-14 cs.CV cs.AI 81%

Difference Vector Equalization for Robust Fine-tuning of Vision-Language Models

Satoshi Suzuki, Shin'ya Yamaguchi, Shoichiro Takeda, Taiga Yamane, Naoki Makishima, Naotaka Kawata, Mana Ihori, Tomohiro Tanaka, Shota Orihashi, Ryo Masumura

专题命中 VLM训练与架构 :vision-language model(title,abstract);分类 cs.CV、cs.AI

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10098 2025-11-14 cs.CV 79%

MTAttack: Multi-Target Backdoor Attacks against Large Vision-Language Models

Zihan Wang, Guansong Pang, Wenjun Miao, Jin Zheng, Xiao Bai

专题命中 VLM训练与架构 :vision-language model(title);visual language model(abstract);分类 cs.CV

Comments AAAI2026, with supplementary material

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09883 2025-11-14 cs.CV 79%

HCC-3D: Hierarchical Compensatory Compression for 98% 3D Token Reduction in Vision-Language Models

Liheng Zhang, Jin Wang, Hui Li, Bingfeng Zhang, Weifeng Liu

专题命中 VLM训练与架构 :vision-language model(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏