arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-07-28 至 2025-07-28 共收录 7 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7 篇

2507.19262 2025-07-28 cs.CV 77%

OVFact: Measuring and Improving Open-Vocabulary Factuality for Long Caption Models

Monika Wysoczańska, Shyamal Buch, Anurag Arnab, Cordelia Schmid

机构 * Google DeepMind(谷歌DeepMind) Warsaw University of Technology(华沙技术大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10568 2025-07-28 cs.CV 70%

AgMMU: A Comprehensive Agricultural Multimodal Understanding Benchmark

Aruna Gauba, Irene Pi, Yunze Man, Ziqi Pang, Vikram S. Adve, Yu-Xiong Wang

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Rice University(Rice大学) Carnegie Mellon University(卡内基梅隆大学) AIFARMS Center for Digital Agriculture at UIUC(伊利诺伊大学厄巴纳-香槟分校数字农业中心)

专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract);分类 cs.CV

Comments Project Website: https://agmmu.github.io/ Huggingface: https://huggingface.co/datasets/AgMMU/AgMMU_v1/

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.07783 2025-07-28 cs.CV cs.CL 70%

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding

Zhaokai Wang, Xizhou Zhu, Xue Yang, Gen Luo, Hao Li, Changyao Tian, Wenhan Dou, Junqi Ge, Lewei Lu, Yu Qiao, Jifeng Dai

机构 * Shanghai Jiao Tong University(上海交通大学) Tsinghua University(清华大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) The Chinese University of Hong Kong(香港中文大学) Sensetime

专题命中 视觉定位与Grounding :LLaVA(abstract);multimodal large language model(abstract);分类 cs.CV

Journal ref TPAMI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05211 2025-07-28 cs.CV cs.AI 62%

All in One: Visual-Description-Guided Unified Point Cloud Segmentation

Zongyan Han, Mohamed El Amine Boudjoghra, Jiahua Dong, Jinhong Wang, Rao Muhammad Anwer

机构 * Mohamed Bin Zayed University of Artificial Intelligence(莫兹哈德大学人工智能大学) Technical University of Munich(慕尼黑技术大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI

Comments Accepted by ICCV2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19370 2025-07-28 cs.CV 57%

BEV-LLM: Leveraging Multimodal BEV Maps for Scene Captioning in Autonomous Driving

Felix Brandstaetter, Erik Schuetz, Katharina Winter, Fabian Flohr

机构 * Intelligent Vehicles Lab (IVL) Munich University of Applied Sciences(智能车辆实验室(IVL)慕尼黑应用科学大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19359 2025-07-28 cs.CV 57%

SemGes: Semantics-aware Co-Speech Gesture Generation using Semantic Coherence and Relevance Learning

Lanmiao Liu, Esam Ghaleb, Aslı Özyürek, Zerrin Yumak

机构 * Max Planck Institute for Psycholinguistics(马克斯·普朗克心理学研究所) Donders Institute for Brain Cognition and Behaviour(多纳尔斯脑认知与行为研究所) Utrecht University(乌得勒支大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments Accepted to IEEE/CVF International Conference on Computer Vision (ICCV) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19213 2025-07-28 cs.CV 57%

PRE-MAP: Personalized Reinforced Eye-tracking Multimodal LLM for High-Resolution Multi-Attribute Point Prediction

Hanbing Wu, Ping Jiang, Anyang Su, Chenxu Zhao, Tianyu Fu, Minghui Wu, Beiping Tan, Huiying Li

机构 * Jilin University(吉林大学) Peking University(北京大学) Mininglamp Technology(Mininglamp科技)

专题命中 视觉定位与Grounding :MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏