arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-09-22 至 2025-09-22 共收录 8 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 8 篇

2509.15532 2025-09-22 cs.CV cs.AI 81%

GUI-ARP: Enhancing Grounding with Adaptive Region Perception for GUI Agents

Xianhang Ye, Yiqing Li, Wei Dai, Miancan Liu, Ziyuan Chen, Zhangye Han, Hongbo Min, Jinkui Ren, Xiantao Zhang, Wen Yang, Zhi Jin

机构 * Wuhan University(武汉大学) Sun Yat-sen University(中山大学) Alibaba Group(阿里巴巴集团) East China Normal University(华东师范大学) University of Electronic Science and Technology of China(电子科技大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15243 2025-09-22 cs.CV 79%

Multi-Modal Interpretability for Enhanced Localization in Vision-Language Models

Muhammad Imran, Yugyung Lee

机构 * Computer Science, School of Science and Engineering, University of Missouri - Kansas City(计算机科学系,科学与工程学院,密苏里大学-堪萨斯城分校)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

Comments 8 pages, 6 figures, 3 tables

Journal ref Non-Archival track - The First Workshop on Multimodal Knowledge and Language Modeling IJCAI 2025 Workshop, August 16, 2025 IJCAI 2025 Workshop, August 16, 2025 Room 516B, Palais des congrès, Montreal, Canada

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15837 2025-09-22 cs.CL 78%

The Curious Case of Visual Grounding: Different Effects for Speech- and Text-based Language Encoders

Adrian Sauter, Willem Zuidema, Marianne de Heer Kloots

机构 * Institute for Logic, Language and Computation(逻辑、语言与计算研究所) University of Amsterdam(阿姆斯特丹大学)

专题命中 视觉定位与Grounding :grounding(title,abstract)

Comments 5 pages, 3 figures, Submitted to ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.13794 2025-09-22 cs.CV cs.AI 73%

LED: LLM Enhanced Open-Vocabulary Object Detection without Human Curated Data Generation

Yang Zhou, Shiyu Zhao, Yuxiao Chen, Zhenting Wang, Can Jin, Dimitris N. Metaxas

机构 * Rutgers University(新泽西罗格斯大学)

专题命中 视觉定位与Grounding :grounding(abstract);MLLM(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15476 2025-09-22 cs.CL cs.MM 71%

Evaluating Multimodal Large Language Models on Spoken Sarcasm Understanding

Zhu Li, Xiyuan Gao, Yuqing Zhang, Shekhar Nayak, Matt Coler

机构 * University of Groningen, The Netherlands(Groningen大学,荷兰)

专题命中 视觉定位与Grounding :multimodal large language model(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14967 2025-09-22 cs.RO cs.HC 67%

Affordance-Based Disambiguation of Surgical Instructions for Collaborative Robot-Assisted Surgery

Ana Davila, Jacinto Colan, Yasuhisa Hasegawa

机构 * Nagoya University, Japan(名古屋大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract)

Comments To be presented at the 1st Workshop on Intelligent Cobodied Assistance and Robotic Empowerment (iCARE). 2025 Conference on Robot Learning (CoRL)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14312 2025-09-22 cs.CV 57%

CLIPTTA: Robust Contrastive Vision-Language Test-Time Adaptation

Marc Lafon, Gustavo Adolfo Vargas Hakim, Clément Rambour, Christian Desrosier, Nicolas Thome

机构 * Conservatoire National des Arts et Métiers(法国国家艺术与工艺学院) Sorbonne Université(索邦大学) ETS Montreal(蒙特利尔ETS) Institut universitaire de France(法国国家科学研究中心)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Journal ref 39th Conference on Neural Information Processing Systems, NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01029 2025-09-22 cs.CY cs.AI cs.DB cs.HC 57%

Who is Responsible When AI Fails? Mapping Causes, Entities, and Consequences of AI Privacy and Ethical Incidents

Hilda Hadan, Reza Hadi Mogavi, Leah Zhang-Kennedy, Lennart E. Nacke

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments 63 pages, 7 tables, 7 figures

Journal ref International Journal of Human-Computer Interaction (2025): 1-45

详情

展开后加载摘要…

URL PDF HTML 收藏