arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-23 至 2025-10-23 共收录 5 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 5 篇

2505.24625 2025-10-23 cs.CV cs.AI 73%

Learning from Videos for 3D World: Enhancing MLLMs with 3D Vision Geometry Priors

Duo Zheng, Shijia Huang, Yanyang Li, Liwei Wang

机构 * The Chinese University of Hong Kong(香港中文大学)

专题命中 视频多模态 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16708 2025-10-23 cs.CL cs.AI 62%

Natural Language Processing for Cardiology: A Narrative Review

Kailai Yang, Yan Leng, Xin Zhang, Tianlin Zhang, Paul Thompson, Bernard Keavney, Maciej Tomaszewski, Sophia Ananiadou

专题命中 视频多模态 :multimodal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19574 2025-10-23 cs.CV cs.CR 57%

Can You Trust What You See? Alpha Channel No-Box Attacks on Video Object Detection

Ariana Yi, Ce Zhou, Liyang Xiao, Qiben Yan

机构 * Mission San Jose High School(Mission San Jose 高中) Missouri University of Science and Technology(密苏里科学与技术大学) Michigan State University(密歇根州立大学)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19560 2025-10-23 cs.CV 57%

HAD: Hierarchical Asymmetric Distillation to Bridge Spatio-Temporal Gaps in Event-Based Object Tracking

Yao Deng, Xian Zhong, Wenxuan Liu, Zhaofei Yu, Jingling Yuan, Tiejun Huang

机构 * Sanya Science and Education Innovation Park, Wuhan University of Technology(武汉理工大学三亚科学教育创新园) Hubei Key Laboratory of Transportation Internet of Things, School of Computer Science and Artificial Intelligence, Wuhan University of Technology(湖北省交通运输物联网重点实验室,计算机科学与人工智能学院,武汉理工大学) State Key Laboratory for Multimedia Information Processing, Peking University(多媒体信息处理国家重点实验室,北京大学)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.21776 2025-10-23 cs.CV 57%

Video-R1: Reinforcing Video Reasoning in MLLMs

Kaituo Feng, Kaixiong Gong, Bohao Li, Zonghao Guo, Yibing Wang, Tianshuo Peng, Junfei Wu, Xiaoying Zhang, Benyou Wang, Xiangyu Yue

机构 * CUHK MMLab(香港中文大学多模态实验室) CUHK (SZ)(香港中文大学(深圳)) Tsinghua University(清华大学) UCAS(中国科学院大学) CUHK HCCL(香港中文大学高性能计算实验室)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments NeurIPS 2025, Project page: https://github.com/tulerfeng/Video-R1

详情

展开后加载摘要…

URL PDF HTML 收藏