arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-08-14 至 2025-08-14 共收录 8 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 8 篇

2507.00669 2025-08-14 cs.LG cs.AI cs.CV cs.RO 82%

Audio-3DVG: Unified Audio -- Point Cloud Fusion for 3D Visual Grounding

Duc Cao-Dinh, Khai Le-Duc, Anh Dao, Bach Phan Tat, Chris Ngo, Duy M. H. Nguyen, Nguyen X. Khanh, Thanh Nguyen-Tang

机构 * Hanyang University(汉阳大学) University of Toronto(多伦多大学) University Health Network(大学健康网络) Knovel Engineering Lab(Knovel工程实验室) Michigan State University(密歇根州立大学) KU Leuven(鲁汶大学) German Research Center for Artificial Intelligence (DFKI)(德国人工智能研究中心) Max Planck Research School for Intelligent Systems (IMPRS-IS)(马克斯·普朗克智能系统研究学校) University of Stuttgart(斯图加特大学) UC Berkeley(伯克利大学) Johns Hopkins University(约翰霍普金斯大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.AI、cs.LG

Comments Preprint, 51 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06564 2025-08-14 cs.CV 74%

Grounding Emotion Recognition with Visual Prototypes: VEGA -- Revisiting CLIP in MERC

Guanyu Hu, Dimitrios Kollias, Xinyu Yang

机构 * Xi'an Jiaotong University(西安交通大学) Queen Mary University of London(伦敦女王玛丽大学) Center for Multimodal AI(多模态人工智能中心) Digital Environment Research Institute(数字环境研究院)

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV

Comments accepted for publication at ACM Multimedia (ACM MM) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09715 2025-08-14 cs.CV cs.LG 62%

NEURAL: Attention-Guided Pruning for Unified Multimodal Resource-Constrained Clinical Evaluation

Devvrat Joshi, Islem Rekik

机构 * BASIRA Lab, Imperial-X (I-X) and Department of Computing, Imperial College London(BASIRA实验室、Imperial-X(I-X)及帝国理工学院计算机系)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.14153 2025-08-14 cs.CV cs.AI cs.CL 62%

Explaining Caption-Image Interactions in CLIP Models with Second-Order Attributions

Lucas Möller, Pascal Tilli, Ngoc Thang Vu, Sebastian Padó

机构 * University of Stuttgart(斯图加特大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI

Comments Accepted at Transactions on Machine Learning Research (TMLR)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18923 2025-08-14 cs.CV 57%

Video SimpleQA: Towards Factuality Evaluation in Large Video Language Models

Meng Cao, Pengfei Hu, Yingyao Wang, Jihao Gu, Haoran Tang, Haoze Zhao, Chen Wang, Jiahua Dong, Wangbo Yu, Ge Zhang, Jun Song, Xiang Li, Bo Zheng, Ian Reid, Xiaodan Liang

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09566 2025-08-14 cs.CV 57%

A Chain of Diagnosis Framework for Accurate and Explainable Radiology Report Generation

Haibo Jin, Haoxuan Che, Sunan He, Hao Chen

机构 * Department of Computer Science and Engineering, Hong Kong University of Science and Technology(计算机科学与工程系,香港科学理工大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments Accepted to IEEE TMI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09293 2025-08-14 cs.CY cs.AI 57%

Ethical Medical Image Synthesis

Weina Jin, Ashish Sinha, Kumar Abhishek, Ghassan Hamarneh

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09245 2025-08-14 cs.CV 57%

Beyond Blanket Masking: Examining Granularity for Privacy Protection in Images Captured by Blind and Low Vision Users

Jeffri Murrugarra-LLerena, Haoran Niu, K. Suzanne Barber, Hal Daumé, Yang Trista Cao, Paola Cascante-Bonilla

机构 * Stony Brook University(石溪大学) University of Texas at Austin(德克萨斯大学奥斯汀分校) University of Maryland(马里兰大学)

专题命中 视觉定位与Grounding :visual language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏