arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-09-16 至 2025-09-16 共收录 9 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 9 篇

2509.10345 2025-09-16 cs.CV cs.AI 88%

Towards Understanding Visual Grounding in Visual Language Models

Georgios Pantazopoulos, Eda B. Özyiğit

机构 * The Alan Turing Institute(艾伦·图灵研究所) Heriot-Watt University(赫瑞-沃德大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);visual language model(title);vision language model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.02573 2025-09-16 cs.CV 83%

Remote Sensing SpatioTemporal Vision-Language Models: A Comprehensive Survey

Chenyang Liu, Jiafan Zhang, Keyan Chen, Man Wang, Zhengxia Zou, Zhenwei Shi

机构 * Department of Aerospace Intelligent Science and Technology, School of Astronautics, Beihang University(航空航天智能科学与技术系,航天学院,北航) Key Laboratory of Spacecraft Design Optimization and Dynamic Simulation Technologies, Ministry of Education, China(航天器设计优化与动态仿真技术重点实验室,教育部,中国) College of Computer Science, Inner Mongolia University(计算机科学学院,内蒙古大学) Shen Yuan Honors College of Beihang University(盛元荣誉学院,北航)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);grounding(abstract);分类 cs.CV

Comments Published in IEEE Geoscience and Remote Sensing Magazine

Journal ref IEEE Geoscience and Remote Sensing Magazine, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11866 2025-09-16 cs.CV 79%

Dr.V: A Hierarchical Perception-Temporal-Cognition Framework to Diagnose Video Hallucination by Fine-grained Spatial-Temporal Grounding

Meng Luo, Shengqiong Wu, Liqiang Jing, Tianjie Ju, Li Zheng, Jinxiang Lai, Tianlong Wu, Xinya Du, Jian Li, Siyuan Yan, Jiebo Luo, William Yang Wang, Hao Fei, Mong-Li Lee, Wynne Hsu

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments 25 pages, 16 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.10424 2025-09-16 cs.CV cs.AI 79%

What is the Visual Cognition Gap between Humans and Multimodal LLMs?

Xu Cao, Yifan Shen, Bolin Lai, Wenqian Ye, Yunsheng Ma, Joerg Heintz, Jintai Chen, Meihuan Huang, Jianguo Cao, Aidong Zhang, James M. Rehg

机构 * Department of Computer Science, University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校计算机科学系) College of Computing, Georgia Institute of Technology(佐治亚理工学院计算机学院) Department of Computer Science, University of Virginia(弗吉尼亚大学计算机科学系) Digital Twin Lab, Purdue University(普渡大学数字孪生实验室) HKUST (Guangzhou)(香港科技大学(广州)) Department of Rehabilitation Medicine, Shenzhen Children’s Hospital(深圳儿童医院康复医学系)

专题命中 视觉定位与Grounding :vision language model(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments COLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12145 2025-09-16 cs.CV 74%

Open-ended Hierarchical Streaming Video Understanding with Vision Language Models

Hyolim Kang, Yunsu Park, Youngbeom Yoo, Yeeun Choi, Seon Joo Kim

机构 * Yonsei University(延世大学)

专题命中 视觉定位与Grounding :vision language model(title);分类 cs.CV

Comments 17 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11840 2025-09-16 cs.CV 57%

Synthetic Captions for Open-Vocabulary Zero-Shot Segmentation

Tim Lebailly, Vijay Veerabadran, Satwik Kottur, Karl Ridgeway, Michael Louis Iuzzolino

机构 * Meta KU Leuven(鲁汶大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments ICCV 2025 CDEL Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11714 2025-09-16 eess.IV cs.LG 57%

EMeRALDS: Electronic Medical Record Driven Automated Lung Nodule Detection and Classification in Thoracic CT Images

Hafza Eman, Furqan Shaukat, Muhammad Hamza Zafar, Syed Muhammad Anwar

机构 * Faculty of Electrical and Electronics Engineering, University of Engineering(电气电子工程学院,工程大学) Department of Engineering Sciences, University of Agder(工程科学系,阿格德大学) Sheikh Zayed Institute for Pediatric Surgical Innovation, Children’s National Hospital(谢赫扎耶德小儿外科创新研究所,儿童医院) School of Medicine and Health Sciences, George Washington University(医学与健康科学学院,乔治华盛顿大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11336 2025-09-16 cs.AI 57%

The power of dynamic causality in observer-based design for soft sensor applications

William Farlessyost, Sebastian Oberst, Shweta Singh

机构 * organization= Agricultural \& Biological Engineering, Purdue University , country= USA organization= Environmental \& Ecological Engineering, Purdue University , country= USA organization= Davidson School of Chemical Engineering, Purdue University , country= USA organization= Mechanical \& Mechatronic Engineering, University of Technology Sydney (UTS) , country= Australia

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10478 2025-09-16 cs.NI cs.LG cs.SY eess.SY 57%

The LLM as a Network Operator: A Vision for Generative AI in the 6G Radio Access Network

Oluwaseyi Giwa, Michael Adewole, Tobi Awodumila, Pelumi Aderinto

机构 * African Institute for Mathematical Sciences(非洲数学科学研究所)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

Comments Submitted to Workshop on AI and ML for Next-Generation Wireless Communications and Networking, NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏