arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-11-13 至 2025-11-13 共收录 33 信号源:cs.CV, cs.AI, cs.LG

1. 视觉问答 3 篇

2506.00942 2025-11-13 cs.CL cs.AI eess.SP 83%

anyECG-chat: A Generalist ECG-MLLM for Flexible ECG Input and Multi-Task Understanding

Haitao Li, Ziyu Li, Yiheng Mao, Ziyi Liu, Zhoujian Sun, Zhengxing Huang

专题命中 视觉问答 :MLLM(title,abstract);multimodal large language model(abstract);分类 cs.AI

Comments AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09058 2025-11-13 cs.CV 79%

VietMEAgent: Culturally-Aware Few-Shot Multimodal Explanation for Vietnamese Visual Question Answering

Hai-Dang Nguyen, Minh-Anh Dang, Minh-Tan Le, Minh-Tuan Le

机构 * Faculty Of Information Technology, VNU University of Engineering and Technology(信息科技学院,越南工程与技术大学) IT-BT Convergence Technology Division, Vietnam-Korea Institute of Science and Technology(IT-BT融合技术部,越南-韩国科学技术院) TADI Global Lab, TADI Global Company Limited(TADI全球实验室,TADI全球公司) Faculty of Finance, Banking Academy of Vietnam(金融学院,越南银行学院)

专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV

Comments 7 pages, 3 figures, 3 tables, FAIR 2025 conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09474 2025-11-13 cs.CV 57%

Surgical AI Copilot: Energy-Based Fourier Gradient Low-Rank Adaptation for Surgical LLM Agent Reasoning and Planning

Jiayuan Huang, Runlong He, Danyal Zaman Khan, Evangelos B. Mazomenos, Danail Stoyanov, Hani Marcus, Linzhe Jiang, Matthew J Clarkson, Mobarak I. Hoque

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

Comments 11 pages

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 视觉推理 7 篇

2510.24792 2025-11-13 cs.CV cs.AI 81%

PISA-Bench: The PISA Index as a Multilingual and Multimodal Metric for the Evaluation of Vision-Language Models

Patrick Haller, Fabio Barth, Jonas Golde, Georg Rehm, Alan Akbik

专题命中 视觉推理 :vision-language model(title,abstract);分类 cs.CV、cs.AI

Comments 8 pages, 11 tables and figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.18145 2025-11-13 cs.CV 79%

CHOICE: Benchmarking the Remote Sensing Capabilities of Large Vision-Language Models

Xiao An, Jiaxing Sun, Zihan Gui, Wei He

机构 * State Key Lab. LIESMARS, Wuhan University(武汉大学遥感信息与地学实验室国家重点实验室)

专题命中 视觉推理 :vision-language model(title,abstract);分类 cs.CV

Comments Accepted by NeurIPS 2025 Track on Datasets and Benchmarks

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09339 2025-11-13 cs.CL 78%

mmJEE-Eval: A Bilingual Multimodal Benchmark for Evaluating Scientific Reasoning in Vision-Language Models

Arka Mukherjee, Shreya Ghosh

机构 * Kalinga Institute of Industrial Technology (KIIT)(喀里亚理工学院) Indian Institute of Technology (IIT), Bhubaneswar(印度理工学院(班加罗尔))

专题命中 视觉推理 :vision-language model(title,abstract)

Comments Accepted to IJCNLP-AACL Findings 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.16975 2025-11-13 cs.CV 70%

EgoTV: Egocentric Task Verification from Natural Language Task Descriptions

Rishi Hazra, Brian Chen, Akshara Rai, Nitin Kamra, Ruta Desai

机构 * Örebro University(奥雷布罗大学) Meta

专题命中 视觉推理 :vision-language model(abstract);grounding(abstract);分类 cs.CV

Comments Accepted at ICCV 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09067 2025-11-13 cs.CL cs.AI 57%

MM-CRITIC: A Holistic Evaluation of Large Multimodal Models as Multimodal Critique

Gailun Zeng, Ziyang Luo, Hongzhan Lin, Yuchen Tian, Kaixin Li, Ziyang Gong, Jianxiong Guo, Jing Ma

机构 * Hong Kong Baptist University(香港 Baptist 大学) Beijing Normal-Hong Kong Baptist University(北京师范大学-香港 Baptist 大学) National University of Singapore(新加坡国立大学) Beijing Normal University(北京师范大学) Shanghai Jiao Tong University(上海交通大学)

专题命中 视觉推理 :visual reasoning(abstract);分类 cs.AI

Comments 28 pages, 14 figures, 19 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23990 2025-11-13 cs.AI 57%

Multi-RAG: A Multimodal Retrieval-Augmented Generation System for Adaptive Video Understanding

Mingyang Mao, Mariela M. Perez-Cabarcas, Utteja Kallakuri, Nicholas R. Waytowich, Xiaomin Lin, Tinoosh Mohsenin

机构 * Johns Hopkins Whiting School of Engineering(约翰霍普金斯大学惠廷工程学院) DEVCOM Army Research Laboratory(国防部陆军研究实验室)

专题命中 视觉推理 :vision-language model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.17582 2025-11-13 cs.HC cs.GR cs.MM 50%

Crafting Dynamic Virtual Activities with Advanced Multimodal Models

Changyang Li, Qingan Yan, Minyoung Kim, Zhan Li, Yi Xu, Lap-Fai Yu

专题命中 视觉推理 :multimodal large language model(abstract)

Journal ref C. Li, Q. Yan, M. Kim, Z. Li, Y. Xu and L. -F. Yu, "Crafting Dynamic Virtual Activities with Advanced Multimodal Models," 2025 IEEE International Symposium on Mixed and Augmented Reality (ISMAR), pp. 120-130

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视觉定位与Grounding 6 篇

2509.10837 2025-11-13 cs.AI 79%

Exploring the Paradigm Shift from Grounding to Skolemization for Complex Query Answering on Knowledge Graphs

Yuyin Lu, Hegang Chen, Shanrui Xie, Yanghui Rao, Haoran Xie, Fu Lee Wang, Qing Li

机构 * School of Computer Science and Engineering, Sun Yat-sen University(中山大学计算机科学与工程学院) School of Data Science, Lingnan University(岭南大学数据科学学院) School of Science and Technology, Hong Kong Metropolitan University(香港理工大学科技学院) Department of Computing, The Hong Kong Polytechnic University(香港理工大学计算机系)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10975 2025-11-13 cs.IR cs.CL 75%

ReFineG: Synergizing Small Supervised Models and LLMs for Low-Resource Grounded Multimodal NER

Jielong Tang, Shuang Wang, Zhenxing Wang, Jianxing Yu, Jian Yin

机构 * School of Artificial Intelligence, Sun Yat-sen University(中山大学人工智能学院) Key Laboratory of Sustainable Tourism Smart Assessment Technology, Ministry of Culture and Tourism, Sun Yat-sen University(文化旅游可持续评估技术重点实验室,中华人民共和国文化和旅游部,中山大学) Beijing Normal University(北京师范大学) Institute of Software, Chinese Academy of Sciences(中国科学院软件研究所)

专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);MLLM(abstract)

Comments CCKS 2025 Shared Task Paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08971 2025-11-13 cs.HC cs.CV cs.MM 70%

Plug-and-Play Clarifier: A Zero-Shot Multimodal Framework for Egocentric Intent Disambiguation

Sicheng Yang, Yukai Huang, Weitong Cai, Shitong Sun, You He, Jiankang Deng, Hang Zhang, Jifei Song, Zhensong Zhang

专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract);分类 cs.CV

Comments 16 pages, 9 figures, AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07983 2025-11-13 cs.CV 57%

ChexFract: From General to Specialized -- Enhancing Fracture Description Generation

Nikolay Nechaev, Evgeniia Przhezdzetskaia, Dmitry Umerenkov, Dmitry V. Dylov

机构 * Artificial Intelligence Research Institute (AIRI)(人工智能研究 institute (AIRI))

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments 13 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10391 2025-11-13 cs.AI 57%

LeanRAG: Knowledge-Graph-Based Generation with Semantic Aggregation and Hierarchical Retrieval

Yaoze Zhang, Rong Wu, Pinlong Cai, Xiaoman Wang, Guohang Yan, Song Mao, Ding Wang, Botian Shi

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments Accepted by AAAI-26

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19594 2025-11-13 cs.CL 50%

Understanding and Leveraging the Expert Specialization of Context Faithfulness in Mixture-of-Experts LLMs

Jun Bai, Minghao Tong, Yang Liu, Zixia Jia, Zilong Zheng

机构 * State Key Laboratory of General Artificial Intelligence, BIGAI(通用人工智能国家重点实验室,BIGAI) School of Computer Science, Wuhan University(武汉大学计算机学院)

专题命中 视觉定位与Grounding :grounding(abstract)

Comments EMNLP 2025 Main

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 文档图表理解 1 篇

2411.07722 2025-11-13 cs.AI 70%

Is Cognition Consistent with Perception? Assessing and Mitigating Multimodal Knowledge Conflicts in Document Understanding

Zirui Shao, Feiyu Gao, Zhaoqing Zhu, Chuwei Luo, Hangdi Xing, Zhi Yu, Qi Zheng, Ming Yan, Jiajun Bu

机构 * Zhejiang Key Laboratory of Accessible Perception and Intelligent Systems, Zhejiang University(浙江可感知智能系统重点实验室,浙江大学) Alibaba Group(阿里巴巴集团) Hangzhou High-Tech Zone (Binjiang) Institute of Blockchain and DataSecurity(杭州高新技术区(滨江)区块链与数据安全研究院)

专题命中 文档图表理解 :multimodal large language model(abstract);MLLM(abstract);分类 cs.AI

Comments Accepted at EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

5. GUI与屏幕智能体 4 篇

2511.08942 2025-11-13 cs.RO cs.AI 83%

Think, Remember, Navigate: Zero-Shot Object-Goal Navigation with VLM-Powered Reasoning

Mobin Habibpour, Fatemeh Afghah

机构 * Holcombe Department of Electrical and Computer Engineering(霍尔科姆电气与计算机工程系) Clemson University(克莱姆森大学)

专题命中 GUI与屏幕智能体 :VLM(title,abstract);vision-language model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08978 2025-11-13 cs.MM cs.CV 79%

Spatio-Temporal Data Enhanced Vision-Language Model for Traffic Scene Understanding

Jingtian Ma, Jingyuan Wang, Wayne Xin Zhao, Guoping Liu, Xiang Wen

机构 * School of Computer Science and Engineering, and the MOE Engineering Research Center of Advanced Computer Application Technology, Beihang University(计算机科学与工程学院,以及教育部先进计算机应用技术工程研究中心,北京航空航天大学) School of Computer Science and Engineering, the School of Economics and Management, and the MIIT Key Laboratory of Data Intelligence and Management, Beihang University(计算机科学与工程学院,经济管理学院,以及工信部数据智能与管理重点实验室,北京航空航天大学) Gaoling School of Artificial Intelligence, Renmin University of China(中关村人工智能学院,中国人民大学) DiDi Global Inc.(滴滴出行公司)

专题命中 GUI与屏幕智能体 :vision-language model(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09127 2025-11-13 cs.AI cs.CL cs.CV cs.HC 62%

History-Aware Reasoning for GUI Agents

Ziwei Wang, Leyang Yang, Xiaoxuan Tang, Sheng Zhou, Dajun Chen, Wei Jiang, Yong Li

专题命中 GUI与屏幕智能体 :multimodal large language model(abstract);分类 cs.CV、cs.AI

Comments Paper accepted to AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08892 2025-11-13 cs.AI 57%

Lumine: An Open Recipe for Building Generalist Agents in 3D Open Worlds

Weihao Tan, Xiangyang Li, Yunhao Fang, Heyuan Yao, Shi Yan, Hao Luo, Tenglong Ao, Huihui Li, Hongbin Ren, Bairen Yi, Yujia Qin, Bo An, Libin Liu, Guang Shi

机构 * ByteDance Seed(字节跳动种子) NTU(国立科技大学)

专题命中 GUI与屏幕智能体 :vision-language model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 幻觉与鲁棒性 4 篇

2511.09018 2025-11-13 cs.CV cs.AI 73%

Causally-Grounded Dual-Path Attention Intervention for Object Hallucination Mitigation in LVLMs

Liu Yu, Zhonghao Chen, Ping Kuang, Zhikun Feng, Fan Zhou, Lan Wang, Gillian Dobbie

机构 * University of Auckland(奥克兰大学) China Scholarship Council(中国留学基金委)

专题命中 幻觉与鲁棒性 :vision-language model(abstract);grounding(abstract);分类 cs.CV、cs.AI

Comments 9 pages, published to AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09228 2025-11-13 cs.CV cs.CL 70%

Taming Object Hallucinations with Verified Atomic Confidence Estimation

Jiarui Liu, Weihao Xuan, Zhijing Jin, Mona Diab

机构 * CMU(卡内基梅隆大学) The University of Tokyo(东京大学) The University of Toronto(多伦多大学)

专题命中 幻觉与鲁棒性 :LLaVA(abstract);multimodal large language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09101 2025-11-13 cs.CV 57%

Ultra-Light Test-Time Adaptation for Vision--Language Models

Byunghyun Kim

机构 * Kyungpook National University(庆尚国立大学)

专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.CV

Comments 7 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.15503 2025-11-13 cs.CV 57%

Domain Adaptation from Generated Multi-Weather Images for Unsupervised Maritime Object Classification

Dan Song, Shumeng Huo, Wenhui Li, Lanjun Wang, Chao Xue, An-An Liu

机构 * The School of Electrical and Information Engineering, Tianjin University, China(天津大学电气与信息工程学院) Tiandy Technologies Co., Ltd, Tianjin, China(天津天盾科技有限公司)

专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

7. VLM训练与架构 6 篇

2511.08914 2025-11-13 cs.CV 83%

SPEED-Q: Staged Processing with Enhanced Distillation towards Efficient Low-bit On-device VLM Quantization

Tianyu Guo, Shanwei Zhao, Shiai Zhu, Chenguang Ma

机构 * Tianyu Guo, Shanwei Zhao, Shiai Zhu, Chenguang Ma(作者)

专题命中 VLM训练与架构 :VLM(title,abstract);vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.17417 2025-11-13 cs.CV 83%

Synth-Align: Improving Trustworthiness in Vision-Language Model with Synthetic Preference Data Alignment

Robert Wijaya, Ngoc-Bao Nguyen, Ngai-Man Cheung

专题命中 VLM训练与架构 :vision-language model(title,abstract);LLaVA(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02615 2025-11-13 astro-ph.IM cs.AI 79%

Radio Astronomy in the Era of Vision-Language Models: Prompt Sensitivity and Adaptation

Mariia Drozdova, Erica Lastufka, Vitaliy Kinakh, Taras Holotyak, Daniel Schaerer, Slava Voloshynovskiy

机构 * University of Geneva(日内瓦大学)

专题命中 VLM训练与架构 :vision-language model(title,abstract);分类 cs.AI

Comments Machine Learning and the Physical Sciences Workshop, NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08970 2025-11-13 astro-ph.SR 78%

JW-Flare: Accurate Solar Flare Forecasting Method Based on Multimodal Large Language Models

Mingfu Shao, Hui Wang, Yuyang Li, Jiaben Lin, Jifeng Liu, Baolin Tan, Juan Guo, Yin Zhang, Jing Huang, Jiangtao Su, Yingzi Sun, Haiqing Xu, Jie Chen, Suo Liu, Yuanyong Deng, Liyue Tong, Yang Bai, Cunshi Wang, Kaifan Ji, Yuqing Zhou

专题命中 VLM训练与架构 :multimodal large language model(title,abstract)

Comments 12 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02671 2025-11-13 cs.CV 74%

Raw Data Matters: Enhancing Prompt Tuning by Internal Augmentation on Vision-Language Models

Haoyang Li, Liang Wang, Chao Wang, Siyu Zhou, Jing Jiang, Yan Peng, Guodong Long

专题命中 VLM训练与架构 :vision-language model(title);分类 cs.CV

Comments 16 pages, 6 figures, 15 tables

详情

展开后加载摘要…

URL PDF HTML 收藏