arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-08-29 至 2025-08-29 共收录 25 信号源:cs.CV, cs.AI, cs.LG

1. 视觉问答 1 篇

2508.19724 2025-08-29 cs.CL cs.AI 57%

NLKI: A lightweight Natural Language Knowledge Integration Framework for Improving Small VLMs in Commonsense VQA Tasks

Aritra Dutta, Swapnanil Mukherjee, Deepanway Ghosal, Somak Aditya

机构 * IIT Kharagpur(印度理工学院Kharagpur分校) Ashoka University(阿什oka大学)

专题命中 视觉问答 :vision-language model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 视觉推理 4 篇

2505.15576 2025-08-29 cs.CV cs.LG 81%

Visual Perturbation and Adaptive Hard Negative Contrastive Learning for Compositional Reasoning in Vision-Language Models

Xin Huang, Ruibin Li, Tong Jia, Wei Zheng, Ya Wang

机构 * School of Artificial Intelligence and Software Engineering, Nanyang Normal University, Henan, China(人工智能与软件工程学院,南阳师范学院,河南) Institute for Artificial Intelligence, Peking University, Beijing, China(人工智能研究院,北京大学,北京) Collaborative Innovation Center of Intelligent Explosion-proof Equipment, Henan, China(智能防爆设备协同创新中心,河南)

专题命中 视觉推理 :vision-language model(title,abstract);分类 cs.CV、cs.LG

Comments Accepted at the International Joint Conference on Artificial Intelligence (IJCAI 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20851 2025-08-29 cs.CV 79%

PathMR: Multimodal Visual Reasoning for Interpretable Pathology Diagnosis

Ye Zhang, Yu Zhou, Jingwen Qi, Yongbing Zhang, Simon Puettmann, Finn Wichmann, Larissa Pereira Ferreira, Lara Sichward, Julius Keyl, Sylvia Hartmann, Shuo Zhao, Hongxiao Wang, Xiaowei Xu, Jianxu Chen

机构 * School of Computer Science and Technology, Harbin Institute of Technology(哈尔滨工业大学计算机科学与技术学院) Leibniz-Institut für Analytische Wissenschaften – ISAS – e.V.(莱比锡分析科学研究所(ISAS)) Department of Pathology, The Sixth Affiliated Hospital, Sun Yat-sen University(中山大学第六附属医院病理科部) Institute of Pathology, University Hospital Essen(埃森大学医院病理科研究所) Academy for Multidisciplinary Studies, Capital Normal University(首都师范大学多学科研究学院)

专题命中 视觉推理 :visual reasoning(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20783 2025-08-29 cs.CV cs.AI 73%

Evaluating Compositional Generalisation in VLMs and Diffusion Models

Beth Pearson, Bilal Boulbarss, Michael Wray, Martha Lewis

机构 * University of Bristol(布里斯托大学) University of Amsterdam(阿姆斯特丹大学)

专题命中 视觉推理 :vision-language model(abstract);VLM(abstract);分类 cs.CV、cs.AI

Comments 11 pages including references, 6 figures. Accepted at IWCS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20345 2025-08-29 cs.CV cs.HC 70%

MedFoundationHub: A Lightweight and Secure Toolkit for Deploying Medical Vision Language Foundation Models

Xiao Li, Yanfan Zhu, Ruining Deng, Wei-Qi Wei, Yu Wang, Shilin Zhao, Yaohong Wang, Haichun Yang, Yuankai Huo

机构 * Vanderbilt University(范德比尔特大学) Weill Cornell Medicine(韦尔·科恩医学中心) Vanderbilt University Medical Center(范德比尔特大学医学中心) UT MD Anderson Cancer Center(德克萨斯大学MD安德森癌症中心)

专题命中 视觉推理 :vision-language model(abstract);LLaVA(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视觉定位与Grounding 8 篇

2508.20188 2025-08-29 cs.CV cs.LG 90%

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study

Max Torop, Masih Eskandar, Nicholas Kurtansky, Jinyang Liu, Jochen Weber, Octavia Camps, Veronica Rotemberg, Jennifer Dy, Kivanc Kose

机构 * Northeastern University(东北大学) Memorial Sloan Kettering Cancer Center(纪念斯隆凯特琳癌症中心)

专题命中 视觉定位与Grounding :grounding(title,abstract);multimodal large language model(title,abstract);MLLM(abstract);分类 cs.CV、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20758 2025-08-29 cs.CV cs.AI 88%

SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding

Jiawen Lin, Shiran Bian, Yihang Zhu, Wenbin Tan, Yachao Zhang, Yuan Xie, Yanyun Qu

机构 * School of Informatics, Xiamen University(厦门大学信息学院) School of Computer Science, Nanjing University(南京大学计算机科学学院) School of Computer Science and Technology, East China Normal University(华东师范大学计算机科学与技术学院) Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University(教育部多媒体可信感知与高效计算重点实验室,厦门大学)

专题命中 视觉定位与Grounding :VLM(title,abstract);grounding(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20279 2025-08-29 cs.CV cs.AI cs.CL 86%

How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding

Zhuoran Yu, Yong Jae Lee

机构 * University of Wisconsin–Madison(威斯康星大学麦迪逊分校)

专题命中 视觉定位与Grounding :grounding(title,abstract);LLaVA(abstract);multimodal large language model(abstract);分类 cs.CV、cs.AI

Comments Accepted by COLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20830 2025-08-29 cs.CV 83%

Estimating 2D Keypoints of Surgical Tools Using Vision-Language Models with Low-Rank Adaptation

Krit Duangprom, Tryphon Lambrou, Binod Bhattarai

机构 * University of Aberdeen(阿伯丁大学)

专题命中 视觉定位与Grounding :vision-language model(title);vision language model(abstract);VLM(abstract);分类 cs.CV

Comments Accepted to MICCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.19573 2025-08-29 cs.CV cs.AI 81%

See then Tell: Enhancing Key Information Extraction with Vision Grounding

Shuhang Liu, Zhenrong Zhang, Pengfei Hu, Jiefeng Ma, Jun Du, Qing Wang, Jianshu Zhang, Chenyu Liu

机构 * University of Science and Technology of China(科学技术大学) iFLYTEK Research(iFLYTEK研究院)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01401 2025-08-29 cs.CV 79%

Language-to-Space Programming for Training-Free 3D Visual Grounding

Boyu Mi, Hanqing Wang, Tai Wang, Yilun Chen, Jiangmiao Pang

机构 * Shanghai Jiao Tong University(上海交通大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20976 2025-08-29 cs.SD cs.AI eess.AS 57%

WoW-Bench: Evaluating Fine-Grained Acoustic Perception in Audio-Language Models via Marine Mammal Vocalizations

Jaeyeon Kim, Heeseung Yun, Sang Hoon Woo, Chao-Han Huck Yang, Gunhee Kim

机构 * Carnegie Mellon University(卡内基梅隆大学) Seoul National University(首尔国立大学) NVIDIA(NVIDIA公司)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments Preprint. Project page: https://jaeyeonkim99.github.io/wow_bench/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01970 2025-08-29 cs.LG 57%

Improving Hospital Risk Prediction with Knowledge-Augmented Multimodal EHR Modeling

Rituparna Datta, Jiaming Cui, Zihan Guan, Vishal G. Reddy, Joshua C. Eby, Gregory Madden, Rupesh Silwal, Anil Vullikanti

机构 * Department of Computer Science, University of Virginia(大学计算机科学系) University of Virginia School of Medicine(弗吉尼亚大学医学院) Virginia Polytechnic Institute and State University(弗吉尼亚理工学院和州立大学) Biocomplexity Institute and Initiative, University of Virginia(大学生物复杂性研究所) Division of Infectious Diseases & International Health, University of Virginia School of Medicine(大学感染性疾病与国际卫生分会)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 文档图表理解 2 篇

2508.18984 2025-08-29 cs.CV 70%

Enhancing Document VQA Models via Retrieval-Augmented Generation

Eric López, Artemis Llabrés, Ernest Valveny

机构 * Computer Vision Center, Universitat Autònoma de Barcelona, Spain(巴塞罗那自治大学计算机视觉中心)

专题命中 文档图表理解 :vision-language model(abstract);visual question answering(abstract);分类 cs.CV

Comments Accepted at Workshop on Machine Learning in Document Analysis and Recognition (ICDAR WML 2025), Wuhan, China

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.14316 2025-08-29 cs.CV 70%

T-Stars-Poster: A Framework for Product-Centric Advertising Image Design

Hongyu Chen, Min Zhou, Jing Jiang, Jiale Chen, Yang Lu, Zihang Lin, Bo Xiao, Tiezheng Ge, Bo Zheng

机构 * Alibaba Group(阿里巴巴集团) Beijing University of Posts and Telecommunications(北京邮电大学)

专题命中 文档图表理解 :VLM(abstract);visual language model(abstract);分类 cs.CV

Comments Accepted by CIKM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 幻觉与鲁棒性 2 篇

2508.20227 2025-08-29 cs.CV cs.AI cs.CL cs.LG 82%

A Novel Framework for Automated Explain Vision Model Using Vision-Language Models

Phu-Vinh Nguyen, Tan-Hanh Pham, Chris Ngo, Truong Son Hy

专题命中 幻觉与鲁棒性 :vision-language model(title,abstract);分类 cs.CV、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06989 2025-08-29 cs.CR cs.CV 70%

Probabilistic Modeling of Jailbreak on Multimodal LLMs: From Quantification to Application

Wenzhuo Xu, Zhipeng Wei, Xiongtao Sun, Zonghao Ying, Deyue Zhang, Dongdong Yang, Xiangzheng Zhang, Quanchen Zou

专题命中 幻觉与鲁棒性 :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

6. VLM训练与架构 6 篇

2508.16188 2025-08-29 cs.CL cs.CV cs.MM cs.SD eess.AS 79%

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation

Weiting Tan, Jiachen Lian, Hirofumi Inaguma, Paden Tomasello, Philipp Koehn, Xutai Ma

机构 * Johns Hopkins University(约翰霍普金斯大学) Meta AI Research(Meta AI 研究)

专题命中 VLM训练与架构 :visual language model(title,abstract);分类 cs.CV

Comments EMNLP 2025 (Findings)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20181 2025-08-29 cs.CV cs.AI cs.CL cs.MM 73%

Mitigating Hallucinations in Multimodal LLMs via Object-aware Preference Optimization

Alberto Compagnoni, Davide Caffagni, Nicholas Moratelli, Lorenzo Baraldi, Marcella Cornia, Rita Cucchiara

机构 * University of Modena and Reggio Emilia(摩德纳和雷吉奥艾米利亚大学)

专题命中 VLM训练与架构 :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments BMVC 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.21066 2025-08-29 cs.CV 70%

OneReward: Unified Mask-Guided Image Generation via Multi-Task Human Preference Learning

Yuan Gong, Xionghui Wang, Jie Wu, Shiyin Wang, Yitong Wang, Xinglong Wu

机构 * ByteDance Inc.(字节跳动公司)

专题命中 VLM训练与架构 :vision-language model(abstract);VLM(abstract);分类 cs.CV

Comments project url: https://one-reward.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16201 2025-08-29 cs.CV cs.AI cs.CL 62%

SpecVLM: Enhancing Speculative Decoding of Video LLMs via Verifier-Guided Token Pruning

Yicheng Ji, Jun Zhang, Heming Xia, Jinpeng Chen, Lidan Shou, Gang Chen, Huan Li

机构 * The State Key Laboratory of Blockchain and Data Security(区块链与数据安全国家重点实验室) Hangzhou High-Tech Zone (Binjiang) Institute of Blockchain and Data Security(杭州高新技术区(滨江)区块链与数据安全研究院) Department of Computing, The Hong Kong Polytechnic University(香港理工大学计算机系) School of Computer Science, Beijing University of Posts and Telecommunications(北京邮电大学计算机学院)

专题命中 VLM训练与架构 :LLaVA(abstract);分类 cs.CV、cs.AI

Comments Accepted at EMNLP 2025 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.21044 2025-08-29 cs.CV 57%

MMG-Vid: Maximizing Marginal Gains at Segment-level and Token-level for Efficient Video LLMs

Junpeng Ma, Qizhe Zhang, Ming Lu, Zhibin Wang, Qiang Zhou, Jun Song, Shanghang Zhang

专题命中 VLM训练与架构 :LLaVA(abstract);分类 cs.CV

Comments 10 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.18674 2025-08-29 cs.CV 57%

Image-guided topic modeling for interpretable privacy classification

Alina Elena Baia, Andrea Cavallaro

专题命中 VLM训练与架构 :vision language model(abstract);分类 cs.CV

Comments Paper accepted at the eXCV Workshop at ECCV 2024. Supplementary material included. Code available at https://github.com/idiap/itm

详情

展开后加载摘要…

URL PDF HTML 收藏

7. 其他VLM 2 篇

2505.10583 2025-08-29 cs.CV cs.CL 79%

Relative Drawing Identification Complexity is Invariant to Modality in Vision-Language Models

Diogo Freitas, Brigt Håvardstun, Cèsar Ferri, Darío Garigliotti, Jan Arne Telle, José Hernández-Orallo

机构 * Interactive Technologies Institute and NOVA LINCS Faculty of Exact Sciences and Engineering University of Madeira Portugal(互动技术研究所和NOVA LINCS精确科学与工程学院马德拉大学) Department of Informatics University of Bergen Norway(信息学院卑尔根大学挪威) Valencian Research Institute for Artificial Intelligence Universitat Politècnica de València Spain(瓦伦西亚人工智能研究机构瓦伦西亚理工大学西班牙) Leverhulme Centre for the Future of Intelligence and Valencian Research Institute for Artificial Intelligence Spain(未来智能中心和瓦伦西亚人工智能研究机构西班牙)

专题命中 其他VLM :vision-language model(title,abstract);分类 cs.CV

Comments 54 pages (42 pages of appendix). Accepted for publication at the ECAI 2025 conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20660 2025-08-29 eess.AS cs.SD 50%

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation

Ruifan Deng, Yitian Gong, Qinghui Gao, Luozhijie Jin, Qinyuan Cheng, Zhaoye Fei, Shimin Li, Xipeng Qiu

机构 * Fudan University(复旦大学) Shanghai Innovation Institute(上海创新研究院)

专题命中 其他VLM :multimodal large language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏