arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-07 至 2025-10-07 共收录 5 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 5 篇

2510.03727 2025-10-07 cs.AI cs.CL cs.CV cs.LG 89%

Bridging the Gap Between Multimodal Foundation Models and World Models

Xuehai He

机构 * Computer Science and Engineering University of California, Santa Cruz(计算机科学与工程大学加州大学圣克ruz分校)

专题命中 视频多模态 :multimodal(title,abstract);multimodal foundation model(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments PhD thesis

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.06461 2025-10-07 cs.CV 79%

Interactive Test-Time Adaptation with Reliable Spatial-Temporal Voxels for Multi-Modal Segmentation

Haozhi Cao, Yuecong Xu, Pengyu Yin, Xingyu Ji, Shenghai Yuan, Jianfei Yang, Lihua Xie

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04819 2025-10-07 cs.CV cs.CL 62%

Visual Representations inside the Language Model

Benlin Liu, Amita Kamath, Madeleine Grunde-McLaughlin, Winson Han, Ranjay Krishna

机构 * University of Washington(华盛顿大学) University of California Los Angeles(加州大学洛杉矶分校) Allen Institute for AI(人工智能研究院)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.CL

Comments Accepted to COLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04753 2025-10-07 cs.CV 57%

Beyond Appearance: Transformer-based Person Identification from Conversational Dynamics

Masoumeh Chapariniya, Teodora Vukovic, Sarah Ebling, Volker Dellwo

机构 * Department of Computational Linguistics, University of Zurich, Zurich, Switzerland(计算语言学系,苏黎世大学,苏黎世,瑞士)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11336 2025-10-07 cs.CV 57%

UGC-VideoCaptioner: An Omni UGC Video Detail Caption Model and New Benchmarks

Peiran Wu, Yunze Liu, Zhengdong Zhu, Enmin Zhou, Junxiao Shen

机构 * University of Bristol(布里斯托大学) Memories.ai Research(Memories.ai研究)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏