arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-11-14 至 2025-11-14 共收录 8 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 8 篇

2511.10212 2025-11-14 cs.CV 86%

Next-Frame Feature Prediction for Multimodal Deepfake Detection and Temporal Localization

Ashutosh Anshul, Shreyas Gopal, Deepu Rajan, Eng Siong Chng

机构 * College of Computing and Data Science(计算与数据科学学院) Nanyang Technological University(南洋理工大学)

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);audio-visual(abstract);分类 cs.CV

Comments Under Review, Multimodal Deepfake detection

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10134 2025-11-14 cs.CV 79%

Explicit Temporal-Semantic Modeling for Dense Video Captioning via Context-Aware Cross-Modal Interaction

Mingda Jia, Weiliang Meng, Zenghuang Fu, Yiheng Li, Qi Zeng, Yifan Zhang, Ju Xin, Rongtao Xu, Jiguang Zhang, Xiaopeng Zhang

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

Comments Accepted to AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09773 2025-11-14 cs.LG eess.SP 78%

NeuroLingua: A Language-Inspired Hierarchical Framework for Multimodal Sleep Stage Classification Using EEG and EOG

Mahdi Samaee, Mehran Yazdi, Daniel Massicotte

专题命中 视频多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10334 2025-11-14 cs.CV 70%

Learning to Tell Apart: Weakly Supervised Video Anomaly Detection via Disentangled Semantic Alignment

Wenti Yin, Huaxin Zhang, Xiang Wang, Yuqing Lu, Yicheng Zhang, Bingquan Gong, Jialong Zuo, Li Yu, Changxin Gao, Nong Sang

专题命中 视频多模态 :multimodal(abstract);multimodal foundation model(abstract);分类 cs.CV

Comments Accepted to AAAI 2026. Code is available at https://github.com/lessiYin/DSANet

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16781 2025-11-14 cs.CV cs.AI cs.CL 67%

Xiaoice: Training-Free Video Understanding via Self-Supervised Spatio-Temporal Clustering of Semantic Features

Shihao Ji, Zihui Song

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments This paper is being withdrawn because we have identified a significant error in the implementation of our self-supervised clustering approach. Specifically, our feature aggregation step inadvertently leaked temporal information across frames, which violates the core assumption of our training-free method. We sincerely apologize to the research community

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25238 2025-11-14 cs.CV 57%

VADB: A Large-Scale Video Aesthetic Database with Professional and Multi-Dimensional Annotations

Qianqian Qiao, DanDan Zheng, Yihang Bo, Bao Peng, Heng Huang, Longteng Jiang, Huaye Wang, Jingdong Chen, Jun Zhou, Xin Jin

机构 * Nanjing University(南京大学) Huazhong University of Science and Technology(华中科技大学) Beijing Film Academy(北京电影学院) University of Science and Technology of China(中国科学技术大学) Beijing Electronic Science and Technology Institute(北京电子科技学院) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Beijing Institute for General Artificial Intelligence(北京通用人工智能研究院)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09870 2025-11-14 cs.CV 57%

SAM-DAQ: Segment Anything Model with Depth-guided Adaptive Queries for RGB-D Video Salient Object Detection

Jia Lin, Xiaofei Zhou, Jiyuan Liu, Runmin Cong, Guodao Zhang, Zhi Liu, Jiyong Zhang

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments Accepted to 40th AAAI Conference on Artificial Intelligence (AAAI 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04369 2025-11-14 cs.CV 57%

TSPO: Temporal Sampling Policy Optimization for Long-form Video Language Understanding

Canhui Tang, Zifan Han, Hongbo Sun, Sanping Zhou, Xuchong Zhang, Xin Wei, Ye Yuan, Huayu Zhang, Jinglin Xu, Hao Sun

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏