arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-10-30 至 2025-10-30 共收录 9 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 9 篇

2510.25032 2025-10-30 cs.CV cs.AI 84%

Efficient License Plate Recognition via Pseudo-Labeled Supervision with Grounding DINO and YOLOv8

Zahra Ebrahimi Vargoorani, Amir Mohammad Ghoreyshi, Ching Yee Suen

专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract);分类 cs.CV、cs.AI

Comments 6 pages, 8 figures. Presented at 2025 IEEE International Workshop on Machine Learning for Signal Processing (MLSP), August 31 - September 3, 2025, Istanbul, Turkey

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25616 2025-10-30 cs.LG cs.AI cs.RO 73%

Don't Blind Your VLA: Aligning Visual Representations for OOD Generalization

Nikita Kachaev, Mikhail Kolosov, Daniil Zelezetsky, Alexey K. Kovalev, Aleksandr I. Panov

机构 * Cognitive AI Lab(认知人工智能实验室) Cognitive AI Lab, IAI MIPT(认知人工智能实验室,IAI MIPT)

专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract);分类 cs.AI、cs.LG

Comments 13 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25368 2025-10-30 cs.LG cs.AI cs.NE 62%

Position: Biology is the Challenge Physics-Informed ML Needs to Evolve

Julien Martinelli

机构 * ELLIS Institute Finland(芬兰ELLIS研究所) Department of Computer Science, Aalto University(艾尔沃斯大学计算机科学系)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25051 2025-10-30 cs.CV cs.LG 62%

Breast Cancer VLMs: Clinically Practical Vision-Language Train-Inference Models

Shunjie-Fabian Zheng, Hyeonjun Lee, Thijs Kooi, Ali Diba

机构 * Department of Medicine I, LMU University Hospital, LMU Munich, Germany(慕尼黑大学医学院第一医学部,LMU慕尼黑大学医院) Lunit Inc.(Lunit公司)

专题命中 视觉定位与Grounding :VLM(abstract);分类 cs.CV、cs.LG

Comments Accepted to Computer Vision for Automated Medical Diagnosis (CVAMD) Workshop at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24737 2025-10-30 eess.SP cs.AI cs.LG 62%

Cardi-GPT: An Expert ECG-Record Processing Chatbot

Koustav Mallick, Neel Singh, Mohammedreza Hajiarbabi

机构 * Department of Computer Science Purdue University Fort Wayne(计算机科学系 Purdue 大学 Fort Wayne)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

Journal ref SoutheastCon 2025 352-357

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25094 2025-10-30 cs.CV 57%

Visual Diversity and Region-aware Prompt Learning for Zero-shot HOI Detection

Chanhyeong Yang, Taehoon Song, Jihwan Park, Hyunwoo J. Kim

机构 * Korea University(韩国大学) Korea Advanced Institute of Science and Technology(韩国科学技术院)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25087 2025-10-30 cs.CL cs.LG 57%

BioCoref: Benchmarking Biomedical Coreference Resolution with LLMs

Nourah M Salem, Elizabeth White, Michael Bada, Lawrence Hunter

机构 * Computational Bioscience Program University of Colorado Anschutz Medical Campus(科学生物学程序,科罗拉多大学安舒茨医学校区) Department of Pediatrics University of Chicago(儿科学系,芝加哥大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25070 2025-10-30 cs.CV 57%

Vision-Language Integration for Zero-Shot Scene Understanding in Real-World Environments

Manjunath Prasad Holenarasipura Rajiv, B. M. Vidyavathi

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments Preprint under review at IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24965 2025-10-30 cs.NE 50%

Exponential Dynamic Energy Network for High Capacity Sequence Memory

Arjun Karuvally, Pichsinee Lertsaroj, Terrence J. Sejnowski, Hava T. Siegelmann

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏