arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-08-18 至 2025-08-18 共收录 6 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 6 篇

2508.11197 2025-08-18 cs.CL cs.AI cs.LG cs.SI 84%

E-CaTCH: Event-Centric Cross-Modal Attention with Temporal Consistency and Class-Imbalance Handling for Misinformation Detection

Ahmad Mousavi, Yeganeh Abdollahinejad, Roberto Corizzo, Nathalie Japkowicz, Zois Boukouvalas

机构 * Department of Mathematics and Statistics, American University, Washington, DC, USA(数学与统计学系,美国大学,华盛顿特区,美国) Department of Computer Science and Mathematics, Pennsylvania State University, Harrisburg, PA, USA(计算机科学与数学系,宾夕法尼亚州立大学,哈里斯堡,宾夕法尼亚州,美国) Department of Computer Science, American University, Washington, DC, USA(计算机科学系,美国大学,华盛顿特区,美国)

专题命中 视频多模态 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11092 2025-08-18 cs.LG 82%

Predictive Multimodal Modeling of Diagnoses and Treatments in EHR

Cindy Shih-Ting Huang, Clarence Boon Liang Ng, Marek Rei

机构 * Imperial College London(帝国理工学院伦敦分校)

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract)

Comments 10 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10922 2025-08-18 cs.CV 79%

A Survey on Video Temporal Grounding with Multimodal Large Language Model

Jianlong Wu, Wei Liu, Ye Liu, Meng Liu, Liqiang Nie, Zhouchen Lin, Chang Wen Chen

机构 * School of Computer Science and Technology, Harbin Institute of Technology(哈尔滨工业大学计算机科学与技术学院) School of Computer Science and Technology, Shandong Jianzhu University(山东建筑大学计算机科学与技术学院) School of Intelligence Science and Technology, Peking University(北京大学智能科学与技术学院) Department of Computing, The Hong Kong Polytechnic University(香港理工大学计算机系)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments 20 pages,6 figures,survey

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11537 2025-08-18 cs.RO 78%

MultiPark: Multimodal Parking Transformer with Next-Segment Prediction

Han Zheng, Zikang Zhou, Guli Zhang, Zhepei Wang, Kaixuan Wang, Peiliang Li, Shaojie Shen, Ming Yang, Tong Qin

机构 * Shanghai Jiao Tong University(上海交通大学) Zhuoyu Technology, Co., Ltd.(珠海宇科技有限公司) Department of Electronic and Computer Engineering, Hong Kong University of Science and Technology(香港理工大学电子与计算机工程系)

专题命中 视频多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01712 2025-08-18 cs.CV cs.AI 62%

HateClipSeg: A Segment-Level Annotated Dataset for Fine-Grained Hate Video Detection

Han Wang, Zhuoran Wang, Roy Ka-Wei Lee

机构 * Singapore University of Technology and Design(新加坡科技设计大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 6 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07992 2025-08-18 cs.MM 57%

Mining the Social Fabric: Unveiling Communities for Fake News Detection in Short Videos

Haisong Gong, Bolan Su, Xinrong Zhang, Jing Li, Qiang Liu, Shu Wu, Liang Wang

专题命中 视频多模态 :multi-modal(abstract);分类 cs.MM

Comments in submission

详情

展开后加载摘要…

URL PDF HTML 收藏