arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-07-28 至 2025-07-28 共收录 4 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 4 篇

2507.19370 2025-07-28 cs.CV 74%

BEV-LLM: Leveraging Multimodal BEV Maps for Scene Captioning in Autonomous Driving

Felix Brandstaetter, Erik Schuetz, Katharina Winter, Fabian Flohr

机构 * Intelligent Vehicles Lab (IVL) Munich University of Applied Sciences(智能车辆实验室(IVL)慕尼黑应用科学大学)

专题命中 图文多模态 :multimodal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18915 2025-07-28 cs.CL cs.CV 62%

Mining Contextualized Visual Associations from Images for Creativity Understanding

Ananya Sahu, Amith Ananthram, Kathleen McKeown

机构 * Columbia University(哥伦比亚大学)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05211 2025-07-28 cs.CV cs.AI 62%

All in One: Visual-Description-Guided Unified Point Cloud Segmentation

Zongyan Han, Mohamed El Amine Boudjoghra, Jiahua Dong, Jinhong Wang, Rao Muhammad Anwer

机构 * Mohamed Bin Zayed University of Artificial Intelligence(莫兹哈德大学人工智能大学) Technical University of Munich(慕尼黑技术大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Accepted by ICCV2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15265 2025-07-28 cs.CV cs.AI cs.CR 62%

Blind Spot Navigation: Evolutionary Discovery of Sensitive Semantic Concepts for LVLMs

Zihao Pan, Yu Tong, Weibin Wu, Jingyi Wang, Lifeng Chen, Zhe Zhao, Jiajia Wei, Yitong Qiao, Zibin Zheng

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments The paper needs major revisions, so it is being withdrawn

详情

展开后加载摘要…

URL PDF HTML 收藏