arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-13 至 2025-10-13 共收录 5 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 5 篇

2510.09266 2025-10-13 cs.CL 79%

CFVBench: A Comprehensive Video Benchmark for Fine-grained Multimodal Retrieval-Augmented Generation

Kaiwen Wei, Xiao Liu, Jie Zhang, Zijian Wang, Ruida Liu, Yuming Yang, Xin Xiao, Xiao Sun, Haoyang Zeng, Changzai Pan, Yidan Zhang, Jiang Zhong, Peijin Wang, Yingchao Feng

机构 * Chongqing University(重庆大学) Independent Researcher(独立研究者) University of the Chinese Academy of Sciences(中国科学院大学) Aerospace Information Research Institute, Chinese Academy of Sciences(中国科学院航天信息研究所)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06509 2025-10-13 cs.CV 74%

From Captions to Keyframes: KeyScore for Multimodal Frame Scoring and Video-Language Understanding

Shih-Yao Lin, Sibendu Paul, Caren Chen

机构 * Amazon Prime Video(亚马逊Prime视频)

专题命中 视频多模态 :multimodal(title);分类 cs.CV

Comments 10 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.02362 2025-10-13 cs.CV cs.AI 62%

A Comprehensive Survey of Mamba Architectures for Medical Image Analysis: Classification, Segmentation, Restoration and Beyond

Shubhi Bansal, Sreeharish A, Madhava Prasath J, Manikandan S, Sreekanth Madisetty, Mohammad Zia Ur Rehman, Chandravardhan Singh Raghaw, Gaurav Duggal, Nagendra Kumar

机构 * Indian Institute of Technology Indore, India R.M.D. Engineering College, Kavaraipettai, India Jio Platforms Limited, India Birla Institute of Technology \& Science Pilani, India

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09200 2025-10-13 cs.CV cs.AI cs.HC 62%

Towards Safer and Understandable Driver Intention Prediction

Mukilan Karuppasamy, Shankar Gangisetty, Shyam Nandan Rai, Carlo Masone, C V Jawahar

机构 * IIIT Hyderabad(海得拉巴印度理工学院) Politecnico di Torino(托里诺理工学院)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.03991 2025-10-13 cs.CV 57%

Deep Learning for Sports Video Event Detection: Tasks, Datasets, Methods, and Challenges

Hao Xu, Arbind Agrahari Baniya, Sam Well, Mohamed Reda Bouadjenek, Richard Dazeley, Sunil Aryal

机构 * School of Information Technology, Deakin University(德肯大学信息科技学院)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments 28 pages

详情

展开后加载摘要…

URL PDF HTML 收藏