arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4735 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4735 篇

2508.16648 2025-08-26 cs.LG cs.AI physics.flu-dyn 57%

LatentFlow: Cross-Frequency Experimental Flow Reconstruction from Sparse Pressure via Latent Mapping

Junle Liu, Chang Liu, Yanyu Ke, Qiuxiang Huang, Jiachen Zhao, Wenliang Chen, K. T. Tse, Gang Hu

专题命中 视频多模态 :cross-modal(abstract);分类 cs.AI

Comments The paper is submitted to IAAI26. Total 9 pages with 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16126 2025-08-25 cs.IR cs.AI 57%

Spacetime-GR: A Spacetime-Aware Generative Model for Large Scale Online POI Recommendation

Haitao Lin, Zhen Yang, Jiawei Xue, Ziji Zhang, Luzhu Wang, Yikun Gu, Yao Xu, Xin Li

机构 * AMAP, Alibaba Group(阿里集团AMAP)

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15903 2025-08-25 cs.CV 57%

VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos

Kaining Li, Shuwei He, Zihan Xu

机构 * Xidian University(西电大学)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11988 2025-08-22 cs.CV 57%

Exploring Spatial-Temporal Dynamics in Event-based Facial Micro-Expression Analysis

Nicolas Mastropasqua, Ignacio Bugueno-Cordova, Rodrigo Verschae, Daniel Acevedo, Pablo Negri, Maria E. Buemi

机构 * Universidad de Buenos Aires, Facultad de Ciencias Exactas y Naturales(布宜诺斯艾利斯大学,精确科学与自然学院) Institute of Engineering Sciences, Universidad de O’Higgins(工程科学研究所,奥希金斯大学) CONICET-UBA, Instituto de Ciencias de la Computacion (ICC)(CONICET-UBA,计算科学研究所) L3S Research Center, Leibniz Universität Hannover(L3S研究中心,汉诺威莱布尼茨大学)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Journal ref 2025 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW); 2nd Workshop on Neuromorphic Vision (NeVi)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15036 2025-08-22 cs.CR cs.AI 57%

MoEcho: Exploiting Side-Channel Attacks to Compromise User Privacy in Mixture-of-Experts LLMs

Ruyi Ding, Tianhong Xu, Xinyi Shen, Aidong Adam Ding, Yunsi Fei

机构 * Louisiana State University(路易斯安那州立大学) Northeastern University(东北大学) Yale University(耶鲁大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI

Comments This paper will appear in CCS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14609 2025-08-21 cs.CV 57%

AnchorSync: Global Consistency Optimization for Long Video Editing

Zichi Liu, Yinggui Wang, Tao Wei, Chao Ma

机构 * MoE Key Lab of Artificial Intelligence, AI Institute Shanghai Jiao Tong University Shanghai China(人工智能联合实验室,人工智能研究院,上海交通大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments ACM MM 2025; Code is released at https://github.com/VISION-SJTU/AnchorSync

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14442 2025-08-21 cs.HC cs.AI 57%

Detecting Reading-Induced Confusion Using EEG and Eye Tracking

Haojun Zhuang, Dünya Baradari, Nataliya Kosmyna, Arnav Balyan, Constanze Albrecht, Stephanie Chen, Pattie Maes

机构 * University of California, Berkeley(加州大学伯克利分校) MIT Media Lab(MIT媒体实验室) Princeton University(普林斯顿大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07961 2025-08-20 cs.CV 57%

Geo4D: Leveraging Video Generators for Geometric 4D Scene Reconstruction

Zeren Jiang, Chuanxia Zheng, Iro Laina, Diane Larlus, Andrea Vedaldi

机构 * Visual Geometry Group, University of Oxford(视觉几何组,牛津大学) Naver Labs Europe(Naver欧洲实验室)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments 17 pages, 6 figures, ICCV 2025 Highlight, Project page: https://geo4d.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.00838 2025-08-20 cs.CV cs.LG 57%

Spatially-guided Temporal Aggregation for Robust Event-RGB Optical Flow Estimation

Qianang Zhou, Junhui Hou, Meiyi Yang, Yongjian Deng, Youfu Li, Junlin Xiong

机构 * Department of Automation, University of Science and Technology of China(自动化系,中国科学技术大学) Department of Computer Science, City University of Hong Kong(计算机科学系,香港城市大学) Department of Mechanical Engineering, City University of Hong Kong(机械工程系,香港城市大学) College of Computer Science, Beijing University of Technology(计算机科学学院,北京理工大学)

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments 11 pages, 8 figures, under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13205 2025-08-20 cs.CV eess.IV 57%

YOLO11-CR: a Lightweight Convolution-and-Attention Framework for Accurate Fatigue Driving Detection

Zhebin Jin, Ligang Dong

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.24039 2025-08-19 cs.CV cs.HC 57%

Foundation Models for Zero-Shot Segmentation of Scientific Images without AI-Ready Data

Shubhabrata Mukherjee, Jack Lang, Obeen Kwon, Iryna Zenyuk, Valerie Brogden, Adam Weber, Daniela Ushizima

机构 * Lawrence Berkeley National Laboratory(伯克利国家实验室) University of California, Irvine(加州大学尔湾分校) University of California, Berkeley(加州大学伯克利分校) Covalent Metrology(协力计量)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments This paper has been accepted for presentation at the 59th International Conference on Parallel Processing (ICPP 2025), DRAI workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02356 2025-08-19 cs.CV 57%

InterRVOS: Interaction-aware Referring Video Object Segmentation

Woojeong Jin, Seongchan Kim, Jaeho Lee, Seungryong Kim

专题命中 视频多模态 :MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07992 2025-08-18 cs.MM 57%

Mining the Social Fabric: Unveiling Communities for Fake News Detection in Short Videos

Haisong Gong, Bolan Su, Xinrong Zhang, Jing Li, Qiang Liu, Shu Wu, Liang Wang

专题命中 视频多模态 :multi-modal(abstract);分类 cs.MM

Comments in submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10784 2025-08-15 q-bio.NC cs.CV 57%

Insights from the Algonauts 2025 Winners

Paul S. Scotti, Mihir Tripathy

机构 * Medical AI Research Center ( MedARC )(医学人工智能研究中心(MedARC)) Baylor College of Medicine(贝勒医学院)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Perspective piece on Algonauts 2025 Challenge conclusion

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07171 2025-08-15 cs.CV 57%

EventRR: Event Referential Reasoning for Referring Video Object Segmentation

Huihui Xu, Jiashi Lin, Haoyu Chen, Junjun He, Lei Zhu

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) The Hong Kong University of Science and Technology(香港科学与技术大学) Northwestern Polytechnical University(西北工业大学)

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00399 2025-08-15 cs.CV 57%

iSafetyBench: A video-language benchmark for safety in industrial environment

Raiyaan Abdullah, Yogesh Singh Rawat, Shruti Vyas

机构 * University of Central Florida(中央佛罗里达大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted to VISION'25 - ICCV 2025 workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.03096 2025-08-15 cs.CV 57%

Scaling Open-Vocabulary Action Detection

Zhen Hao Sia, Yogesh Singh Rawat

机构 * University of Central Florida(中央佛罗里达大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09857 2025-08-14 cs.CV 57%

OneVAE: Joint Discrete and Continuous Optimization Helps Discrete Video VAE Train Better

Yupeng Zhou, Zhen Li, Ziheng Ouyang, Yuming Chen, Ruoyi Du, Daquan Zhou, Bin Fu, Yihao Liu, Peng Gao, Ming-Ming Cheng, Qibin Hou

机构 * VCIP, School of Computer Science, Nankai University(VCIP,计算机科学学院,南开大学) Shanghai AI Laboratory(上海人工智能实验室) The Chinese University of Hong Kong(香港中文大学) Peking University(北京大学)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18923 2025-08-14 cs.CV 57%

Video SimpleQA: Towards Factuality Evaluation in Large Video Language Models

Meng Cao, Pengfei Hu, Yingyao Wang, Jihao Gu, Haoran Tang, Haoze Zhao, Chen Wang, Jiahua Dong, Wangbo Yu, Ge Zhang, Jun Song, Xiang Li, Bo Zheng, Ian Reid, Xiaodan Liang

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09262 2025-08-14 cs.CV cs.LG 57%

Harnessing Input-Adaptive Inference for Efficient VLN

Dongwoo Kang, Akhil Perincherry, Zachary Coalson, Aiden Gabriel, Stefan Lee, Sanghyun Hong

机构 * Oregon State University(俄勒冈州立大学)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments Accepted to ICCV 2025 [Poster]

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08989 2025-08-13 cs.CV 57%

KFFocus: Highlighting Keyframes for Enhanced Video Understanding

Ming Nie, Chunwei Wang, Hang Xu, Li Zhang

机构 * School of Data Science, Fudan University(复旦大学数据科学学院)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08590 2025-08-13 cs.CV cs.HC 57%

QueryCraft: Transformer-Guided Query Initialization for Enhanced Human-Object Interaction Detection

Yuxiao Wang, Wolin Liang, Yu Lei, Weiying Xue, Nan Zhuang, Qi Liu

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07989 2025-08-12 cs.CV cs.HC 57%

The Escalator Problem: Identifying Implicit Motion Blindness in AI for Accessibility

Xiantao Zhang

机构 * Beihang University(北航大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments 9 pages, 3 figures, 2 tables. Accepted at CV4A11y, ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07626 2025-08-12 cs.CV cs.RO 57%

AR-VRM: Imitating Human Motions for Visual Robot Manipulation with Analogical Reasoning

Dejie Yang, Zijing Zhao, Yang Liu

机构 * Wangxuan Institute of Computer Technology, Peking University(北京大学计算机技术研究院) State Key Laboratory of General Artificial Intelligence, Peking University(北京大学通用人工智能国家重点实验室)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments Accepted by ICCV2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07312 2025-08-12 cs.CV 57%

MobileViCLIP: An Efficient Video-Text Model for Mobile Devices

Min Yang, Zihan Jia, Zhilin Dai, Sheng Guo, Limin Wang

机构 * State Key Laboratory for Novel Software Technology, Nanjing University(新型软件技术国家重点实验室,南京大学) MyBank, Ant Group(蚂蚁集团MyBank) Shanghai AI Lab(上海AI实验室)

专题命中 视频多模态 :image-text(abstract);分类 cs.CV

Comments Accepted by ICCV2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07006 2025-08-12 eess.IV cs.CV 57%

Spatio-Temporal Conditional Diffusion Models for Forecasting Future Multiple Sclerosis Lesion Masks Conditioned on Treatments

Gian Mario Favero, Ge Ya Luo, Nima Fathi, Justin Szeto, Douglas L. Arnold, Brennan Nichyporuk, Chris Pal, Tal Arbel

机构 * McGill University(麦吉尔大学) Mila – Quebec AI Institute(魁北克人工智能研究所)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments Accepted to MICCAI 2025 (LMID Workshop)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12524 2025-08-12 cs.CV cs.HC cs.LG eess.IV 57%

Inference-Time Gaze Refinement for Micro-Expression Recognition: Enhancing Event-Based Eye Tracking with Motion-Aware Post-Processing

Nuwan Bandara, Thivya Kandappu, Archan Misra

机构 * School of Computing(计算学院) Information Systems, Singapore Management University, Singapore(信息系统,新加坡管理大学,新加坡)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted at 4DMR@IJCAI25: International IJCAI Workshop on 1st Challenge and Workshop for 4D Micro-Expression Recognition for Mind Reading, August 29, 2025, Guangzhou, China

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.04592 2025-08-11 cs.LG cs.AI cs.CE cs.IR 57%

CAMEF: Causal-Augmented Multi-Modality Event-Driven Financial Forecasting by Integrating Time Series Patterns and Salient Macroeconomic Announcements

Yang Zhang, Wenbo Yang, Jun Wang, Qiang Ma, Jie Xiong

机构 * Southwestern University of Finance and Economics(西南财经大学) Kyoto Institute of Technology(京都技术大学)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.AI

Comments Accepted in SIGKDD 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.13667 2025-08-11 cs.CV 57%

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation

Fu Rong, Meng Lan, Qian Zhang, Lefei Zhang

机构 * National Engineering Research Center for Multimedia Software, School of Computer Science, Wuhan University(国家多媒体软件工程研究中心,计算机学院,武汉大学) Hong Kong University of Science and Technology(香港科学与技术大学) Horizon Robotics

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.10683 2025-08-11 cs.RO cs.AI 57%

Learning to Initialize Trajectory Optimization for Vision-Based Autonomous Flight in Unknown Environments

Yicheng Chen, Jinjie Li, Wenyuan Qin, Yongzhao Hua, Xiwang Dong, Qingdong Li

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI

Comments Accepted to IROS 2025. Source code available

详情

展开后加载摘要…

URL PDF HTML 收藏