arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-09-04 至 2025-09-04 共收录 3 信号源:cs.CV, cs.AI, cs.LG

1. VLM训练与架构 3 篇

2509.02805 2025-09-04 cs.LG 85%

Challenges in Understanding Modality Conflict in Vision-Language Models

Trang Nguyen, Jackson Michaels, Madalina Fiterau, David Jensen

机构 * Manning College of Information \& Computer Sciences, University of Massachusetts Amherst, Amherst, U.S.

专题命中 VLM训练与架构 :vision-language model(title,abstract);VLM(abstract);LLaVA(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02659 2025-09-04 cs.CV cs.RO 83%

2nd Place Solution for CVPR2024 E2E Challenge: End-to-End Autonomous Driving Using Vision Language Model

Zilong Guo, Yi Luo, Long Sha, Dongxu Wang, Panqu Wang, Chenyang Xu, Yi Yang

机构 * ZERON Shanghai, China(上海零点科技有限公司)

专题命中 VLM训练与架构 :vision language model(title,abstract);VLM(abstract);分类 cs.CV

Comments 2nd place in CVPR 2024 End-to-End Driving at Scale Challenge

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16423 2025-09-04 cs.CV cs.LG 62%

GAEA: A Geolocation Aware Conversational Assistant

Ron Campos, Ashmal Vayani, Parth Parag Kulkarni, Rohit Gupta, Aizan Zafar, Aritra Dutta, Mubarak Shah

机构 * Center for Research in Computer Vision, University of Central Florida(计算机视觉研究中心,中央佛罗里达大学)

专题命中 VLM训练与架构 :LLaVA(abstract);分类 cs.CV、cs.LG

Comments The dataset and code used in this submission is available at: https://ucf-crcv.github.io/GAEA/

详情

展开后加载摘要…

URL PDF HTML 收藏