arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-27 至 2025-10-27 共收录 12 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 12 篇

2510.21445 2025-10-27 cs.CL cs.AI cs.CV cs.LG 82%

REMONI: An Autonomous System Integrating Wearables and Multimodal Large Language Models for Enhanced Remote Health Monitoring

Thanh Cong Ho, Farah Kharrat, Abderrazek Abid, Fakhri Karray

机构 * 2 Department of Electrical Computer Engineering University of Waterloo, Waterloo, ON, Canada N2L 3G1 Email 3 College of Computer Information Sciences Prince Sultan University Email

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Journal ref 2024 IEEE International Symposium on Medical Measurements and Applications (MeMeA)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21406 2025-10-27 cs.CV 79%

MUVR: A Multi-Modal Untrimmed Video Retrieval Benchmark with Multi-Level Visual Correspondence

Yue Feng, Jinwei Hu, Qijia Lu, Jiawei Niu, Li Tan, Shuo Yuan, Ziyi Yan, Yizhen Jia, Qingzhi He, Shiping Ge, Ethan Q. Chen, Wentong Li, Limin Wang, Jie Qin

机构 * MoE Key Laboratory of Brain-Machine Intelligence Technology, College of Artificial Intelligence, Nanjing University of Aeronautics and Astronautics(脑机智能技术MoE实验室,人工智能学院,南京航空航天大学) Nanjing University(南京大学) The Hong Kong Polytechnic University(香港理工大学)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted to NeurIPS 2025 D&B Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03340 2025-10-27 cs.CV 79%

Seeing the Arrow of Time in Large Multimodal Models

Zihui Xue, Mi Luo, Kristen Grauman

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments Accepted by NeurIPS 2025, Project website: https://vision.cs.utexas.edu/projects/SeeAoT

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20952 2025-10-27 cs.LG 78%

LLM-Integrated Bayesian State Space Models for Multimodal Time-Series Forecasting

Sungjun Cho, Changho Shin, Suenggwan Jo, Xinya Yan, Shourjo Aditya Chaudhuri, Frederic Sala

机构 * Department of Computer Sciences University of Wisconsin-Madison(计算机科学系威斯康星大学麦迪逊分校)

专题命中 视频多模态 :multimodal(title,abstract)

Comments 15 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18438 2025-10-27 cs.CR 71%

DeepTx: Real-Time Transaction Risk Analysis via Multi-Modal Features and LLM Reasoning

Yixuan Liu, Xinlei Li, Yi Li

专题命中 视频多模态 :multi-modal(title)

Comments Accepted to ASE'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.03674 2025-10-27 cs.CV cs.AI 62%

Action Quality Assessment via Hierarchical Pose-guided Multi-stage Contrastive Regression

Mengshi Qi, Hao Ye, Jiaxuan Peng, Huadong Ma

机构 * State Key Laboratory of Networking and Switching Technology, Beijing University of Posts and Telecommunications, China(网络与交换技术国家重点实验室,北京邮电大学)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20189 2025-10-27 cs.CV 57%

SPAN: Continuous Modeling of Suspicion Progression for Temporal Intention Localization

Xinyi Hu, Yuran Wang, Ruixu Zhang, Yue Li, Wenxuan Liu, Zheng Wang

机构 * National Engineering Research Center for Multimedia Software, Institute of Artificial Intelligence, School of Computer Science(多媒体软件国家工程研究中心、人工智能研究院、计算机科学学院) Hubei Key Laboratory of Multimedia and Network Communication Engineering(多媒体与网络通信工程湖北省重点实验室) School of Mathematical Sciences, Peking University(北京大学数学科学学院) Tsinghua University(清华大学) School of Computer Science, Peking University(北京大学计算机科学学院) State Key Laboratory for Multimedia Information Processing, Peking University(多媒体信息处理国家重点实验室)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08974 2025-10-27 cs.CV 57%

Text-conditioned State Space Model For Domain-generalized Change Detection Visual Question Answering

Elman Ghazaei, Erchan Aptoula

机构 * Faculty of Engineering and Natural Sciences (VPALab)(工程与自然科学学院(VPALab)) Sabanci University(萨班奇大学)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18812 2025-10-27 cs.CV 57%

SAMA: Towards Multi-Turn Referential Grounded Video Chat with Large Language Models

Ye Sun, Hao Zhang, Henghui Ding, Tiehua Zhang, Xingjun Ma, Yu-Gang Jiang

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21107 2025-10-27 cs.LG cs.AI cs.RO 57%

ESCORT: Efficient Stein-variational and Sliced Consistency-Optimized Temporal Belief Representation for POMDPs

Yunuo Zhang, Baiting Luo, Ayan Mukhopadhyay, Gabor Karsai, Abhishek Dubey

机构 * Vanderbilt University(范德比大学)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.AI

Comments Proceeding of the 39th Conference on Neural Information Processing Systems (NeurIPS'25). Code would be available at https://github.com/scope-lab-vu/ESCORT

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20951 2025-10-27 cs.CV 57%

Generative Point Tracking with Flow Matching

Mattie Tesfaldet, Adam W. Harley, Konstantinos G. Derpanis, Derek Nowrouzezahrai, Christopher Pal

机构 * McGill University(麦吉尔大学) Mila Stanford University(斯坦福大学) York University(约克大学) Polytechnique Montréal(蒙特利尔理工学院)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments Project page: https://mtesfaldet.net/genpt_projpage/

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15745 2025-10-27 eess.IV cs.LG 50%

InfiniPot-V: Memory-Constrained KV Cache Compression for Streaming Video Understanding

Minsoo Kim, Kyuhong Shim, Jungwook Choi, Simyung Chang

专题命中 视频多模态 :multimodal(abstract)

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏