arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-11-14 至 2025-11-14 共收录 2 信号源:cs.CV, cs.AI, cs.LG

1. 其他VLM 2 篇

2409.17806 2025-11-14 cs.LG 79%

Caption, Create, Continue: Continual Learning with Pre-trained Generative Vision-Language Models

Indu Solomon, Aye Phyu Phyu Aung, Uttam Kumar, Senthilnath Jayavelu

机构 * International Institute of Information Technology Bangalore (IIITB), India(国际信息技术研究所(班加罗尔)) Institute for Infocomm Research, Agency for Science, Technology and Research (A*STAR), Singapore(信息与通信研究所(A*STAR))

专题命中 其他VLM :vision-language model(title,abstract);分类 cs.LG

Comments This is the revised and peer-reviewed version of our paper, accepted and published in the Proceedings of the 34th ACM International Conference on Information and Knowledge Management (CIKM 2025)

Journal ref Proc. 34th ACM International Conference on Information and Knowledge Management (CIKM), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26125 2025-11-14 cs.CV cs.AI 62%

WOD-E2E: Waymo Open Dataset for End-to-End Driving in Challenging Long-tail Scenarios

Runsheng Xu, Hubert Lin, Wonseok Jeon, Hao Feng, Yuliang Zou, Liting Sun, John Gorman, Ekaterina Tolstaya, Sarah Tang, Brandyn White, Ben Sapp, Mingxing Tan, Jyh-Jing Hwang, Dragomir Anguelov

机构 * Waymo LLC(Waymo公司)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏