arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-07-30 至 2025-07-30 共收录 8 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 8 篇

2507.21100 2025-07-30 cs.CY cs.AI cs.CV 84%

A Tactical Behaviour Recognition Framework Based on Causal Multimodal Reasoning: A Study on Covert Audio-Video Analysis Combining GAN Structure Enhancement and Phonetic Accent Modelling

Wei Meng

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments This paper introduces a structurally innovative and mathematically rigorous framework for multimodal tactical reasoning, offering a significant advance in causal inference and graph-based threat recognition under noisy conditions

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21945 2025-07-30 cs.CV 83%

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment

Xin Wang, Peng-Jie Li, Yuan-Yuan Shen

机构 * School of Sport Engineering, Beijing Sport University(体育工程学院,北京体育大学)

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted to Applied Soft Computing

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21161 2025-07-30 cs.CV cs.AI cs.LG 81%

Seeing Beyond Frames: Zero-Shot Pedestrian Intention Prediction with Raw Temporal Video and Multimodal Cues

Pallavi Zambare, Venkata Nikhil Thanikella, Ying Liu

机构 * Departmrnt of computer science(计算机科学系) Texas Tech University(得克萨斯科技大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted in IEEE 3rd International Conference on Artificial Intelligence, Blockchain, and Internet of Things (AIBThings 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21649 2025-07-30 cs.CV 79%

The Evolution of Video Anomaly Detection: A Unified Framework from DNN to MLLM

Shibo Gao, Peipei Yang, Haiyang Guo, Yangyang Liu, Yi Chen, Shuai Li, Han Zhu, Jian Xu, Xu-Yao Zhang, Linlin Huang

机构 * Beijing Jiaotong University(北京交通大学) State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences(多模态人工智能系统国家重点实验室,自动化研究所,中国科学院) School of Artificial Intelligence, University of Chinese Academy of Sciences(人工智能学院,中国科学院大学) Zhongguancun Academy, Beijing, China(中关村学院,北京,中国)

专题命中 视频多模态 :MLLM(title);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10541 2025-07-30 cs.IR cs.AI 79%

Multi-Modal Hypergraph Enhanced LLM Learning for Recommendation

Xu Guo, Tong Zhang, Yuanzhi Wang, Chenxu Wang, Fuyun Wang, Xudong Wang, Xiaoya Zhang, Xin Liu, Zhen Cui

机构 * School of Computer Science and Engineering, Nanjing University of Science and Technology(南京理工大学计算机科学与工程学院) Shituoyun (Nanjing) Technology Co., Ltd(石图云(南京)科技有限公司) School of Artificial Intelligence, Beijing Normal University(北京师范大学人工智能学院)

专题命中 视频多模态 :multi-modal(title);multimodal(abstract);分类 cs.AI

Comments 12 pages, 4 figures, submitted to IEEE Transactions on Knowledge and Data Engineering

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20177 2025-07-30 cs.CV cs.MM 73%

Towards Universal Modal Tracking with Online Dense Temporal Token Learning

Yaozong Zheng, Bineng Zhong, Qihua Liang, Shengping Zhang, Guorong Li, Xianxian Li, Rongrong Ji

机构 * Key Laboratory of Education Blockchain and Intelligent Technology, Ministry of Education(教育区块链与智能技术重点实验室,教育部) Guangxi Key Laboratory of Multi-Source Information Mining and Security, Guangxi Normal University(广西多源信息挖掘与安全重点实验室,广西师范大学) School of Computer Science and Technology, Harbin Institute of Technology(哈尔滨工业大学计算机科学与技术学院) University of Chinese Academy of Sciences(中国科学院大学) Media Analytics and Computing Lab, Department of Artificial Intelligence, School of Informatics, Xiamen University(媒体分析与计算实验室,人工智能系,厦门大学)

专题命中 视频多模态 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.MM

Comments arXiv admin note: text overlap with arXiv:2401.01686

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21063 2025-07-30 q-bio.NC cs.CY 71%

Make Silence Speak for Itself: a multi-modal learning analytic approach with neurophysiological data

Mingxuan Gao, Jingjing Chen, Yun Long, Xiaomeng Xu, Yu Zhang

专题命中 视频多模态 :multi-modal(title)

Comments 25 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21971 2025-07-30 cs.CV 57%

EIFNet: Leveraging Event-Image Fusion for Robust Semantic Segmentation

Zhijiang Li, Haoran He

机构 * Beijing University of Posts and Telecommunications(北京邮电大学)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏