arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4757 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4757 篇

2412.17252 2025-07-24 cs.LG math.OC 78%

A Coalition Game for On-demand Multi-modal 3D Automated Delivery System

Farzan Moosavi, Bilal Farooq

机构 * Laboratory of Innovations in Transportation (LiTrans), Toronto Metropolitan University, Toronto, Canada(创新交通实验室(LiTrans),多伦多 Metropolitan 大学,多伦多,加拿大)

专题命中 视频多模态 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14163 2025-07-22 eess.SP cs.LG stat.ML 78%

UniPhyNet: A Unified Network For Multimodal Physiological Raw Signal Classification

Renxiang Qiu, Raghavendra Selvan

机构 * Department of Computer Science, University of Copenhagen(计算机科学系,哥本哈根大学)

专题命中 视频多模态 :multimodal(title,abstract)

Comments Accepted to be presented at the 35th IEEE International Workshop on Machine Learning for Signal Processing (IEEE MLSP 2025). Source code available at https://github.com/HughYau/UniPhyNet

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14146 2025-07-22 eess.SP 78%

Estimating Markers of Driving Stress through Multimodal Physiological Monitoring

Kleanthis Avramidis, Emily Zhou, Tiantian Feng, Hossein Hamidi Shishavan, Frederico Marcolino Quintao Severgnini, Danny J. Lohan, Paul Schmalenberg, Ercan M. Dede, Shrikanth Narayanan

专题命中 视频多模态 :multimodal(title,abstract)

Comments 11 pages, 7 figures, 3 tables. This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06444 2025-07-17 cs.CE 78%

Eyes on the Road, Mind Beyond Vision: Context-Aware Multi-modal Enhanced Risk Anticipation

Jiaxun Zhang, Haicheng Liao, Yumu Xie, Chengyue Wang, Yanchen Guan, Bin Rao, Zhenning Li

专题命中 视频多模态 :multi-modal(title,abstract)

Comments Accepted by ACMMM2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07804 2025-07-11 cs.LG 78%

Deep Survival Analysis in Multimodal Medical Data: A Parametric and Probabilistic Approach with Competing Risks

Alba Garrido, Alejandro Almodóvar, Patricia A. Apellániz, Juan Parras, Santiago Zazo

机构 * Information Processing and Telecommunications Center, ETSI Telecomunicación, Universidad Politécnica de Madrid, Spain(信息处理与电信中心,电信工程学院,马德里理工大学,西班牙)

专题命中 视频多模态 :multimodal(title,abstract)

Comments 29 pages, 9 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18585 2025-06-27 gr-qc 78%

Multimodal signatures of asymptotic (A)dS Kalb-Ramond black holes: Constraints through the shadow, weak deflection angle, and topological photon spheres

Reggie C. Pantig, Ali Övgün

专题命中 视频多模态 :multimodal(title,abstract)

Comments 15 pages, 4 figures

Journal ref Annals of Physics 480 (2025) 170104

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18538 2025-05-27 eess.IV cs.LG 78%

Mind Your Vision: Multimodal Estimation of Refractive Disorders Using Electrooculography and Eye Tracking

Xin Wei, Huakun Liu, Yutaro Hirao, Monica Perusquia-Hernandez, Katsutoshi Masai, Hideaki Uchiyama, Kiyoshi Kiyokawa

机构 * Nara Institute of Science and Technology(奈良科学技術大學) Kyushu University(九州大學)

专题命中 视频多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.18815 2025-05-27 cs.LG 78%

MissionGNN: Hierarchical Multimodal GNN-based Weakly Supervised Video Anomaly Recognition with Mission-Specific Knowledge Graph Generation

Sanggeon Yun, Ryozo Masukawa, Minhyoung Na, Mohsen Imani

专题命中 视频多模态 :multimodal(title,abstract)

Comments Accepted to WACV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11214 2025-05-19 cs.RO 78%

Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions

Wei Zhao, Gongsheng Li, Zhefei Gong, Pengxiang Ding, Han Zhao, Donglin Wang

机构 * Westlake University(西湖大学) Zhejiang University(浙江大学)

专题命中 视频多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.18438 2025-05-08 cs.HC 78%

Adaptive Gen-AI Guidance in Virtual Reality: A Multimodal Exploration of Engagement in Neapolitan Pizza-Making

Ka Hei Carrie Lau, Sema Sen, Philipp Stark, Efe Bozkir, Enkelejda Kasneci

专题命中 视频多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01989 2025-05-06 cs.DS 78%

Exact Set Packing in Multimodal Transportation with Ridesharing System for First/Last Mile

Qian-Ping Gu, Jiajian Leo Liang

专题命中 视频多模态 :multimodal(title,abstract)

Comments 29 pages, 9 tables, 2 figures, and

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10921 2025-04-28 cs.IR 78%

MSCRS: Multi-modal Semantic Graph Prompt Learning Framework for Conversational Recommender Systems

Yibiao Wei, Jie Zou, Weikang Guo, Guoqing Wang, Xing Xu, Yang Yang

专题命中 视频多模态 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.14996 2025-04-08 eess.SY cs.RO cs.SY 78%

EDRF: Enhanced Driving Risk Field Based on Multimodal Trajectory Prediction and Its Applications

Junkai Jiang, Zeyu Han, Yuning Wang, Mengchi Cai, Qingwen Meng, Qing Xu, Jianqiang Wang

专题命中 视频多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.03423 2025-04-07 cs.LG cs.RO 78%

DML-RAM: Deep Multimodal Learning Framework for Robotic Arm Manipulation using Pre-trained Models

Sathish Kumar, Swaroop Damodaran, Naveen Kumar Kuruba, Sumit Jha, Arvind Ramanathan

专题命中 视频多模态 :multimodal(title,abstract)

Comments 7 pages , 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.13683 2025-03-12 cs.RO 78%

PrefMMT: Modeling Human Preferences in Preference-based Reinforcement Learning with Multimodal Transformers

Dezhong Zhao, Ruiqi Wang, Dayoon Suh, Taehyeon Kim, Ziqin Yuan, Byung-Cheol Min, Guohua Chen

专题命中 视频多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.11144 2025-02-19 cs.HC 78%

CrossA11y: Identifying Video Accessibility Issues via Cross-modal Grounding

Xingyu "Bruce" Liu, Ruolin Wang, Dingzeyu Li, Xiang 'Anthony' Chen, Amy Pavel

专题命中 视频多模态 :cross-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.13464 2025-01-24 eess.SP 78%

Deep Multi-modal Neural Receiver for 6G Vehicular Communication

Osama Saleem, Mohammed Alfaqawi, Pierre Merdrignac, Abdelaziz Bensrhair, Soheyb Ribouh

专题命中 视频多模态 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.08486 2025-01-23 cs.HC 78%

Can AI Prompt Humans? Multimodal Agents Prompt Players' Game Actions and Show Consequences to Raise Sustainability Awareness

Qinshi Zhang, Ruoyu Wen, Latisha Besariani Hendra, Zijian Ding, Ray LC

专题命中 视频多模态 :multimodal(title,abstract)

Comments 25 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.00771 2025-01-13 cs.ET 78%

Resistive memory-based zero-shot liquid state machine for multimodal event data learning

Ning Lin, Shaocong Wang, Yi Li, Bo Wang, Shuhui Shi, Yangu He, Woyu Zhang, Yifei Yu, Yue Zhang, Xinyuan Zhang, Kwunhang Wong, Songqi Wang, Xiaoming Chen, Hao Jiang, Xumeng Zhang, Peng Lin, Xiaoxin Xu, Xiaojuan Qi, Zhongrui Wang, Dashan Shang, Qi Liu, Ming Liu

专题命中 视频多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.12410 2024-11-26 cs.LG eess.SP stat.ML 78%

Deep sr-DDL: Deep Structurally Regularized Dynamic Dictionary Learning to Integrate Multimodal and Dynamic Functional Connectomics data for Multidimensional Clinical Characterizations

Niharika Shimona D'Souza, Mary Beth Nebel, Deana Crocetti, Nicholas Wymbs, Joshua Robinson, Stewart H. Mostofsky, Archana Venkataraman

专题命中 视频多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2007.01931 2024-11-25 cs.LG eess.SP stat.ML 78%

A Deep-Generative Hybrid Model to Integrate Multimodal and Dynamic Connectivity for Predicting Spectrum-Level Deficits in Autism

Niharika Shimona D'Souza, Mary Beth Nebel, Deana Crocetti, Nicholas Wymbs, Joshua Robinson, Stewart Mostofsky, Archana Venkataraman

专题命中 视频多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.09577 2024-11-19 cs.HC 78%

SimTube: Generating Simulated Video Comments through Multimodal AI and User Personas

Yu-Kai Hung, Yun-Chien Huang, Ting-Yu Su, Yen-Ting Lin, Lung-Pan Cheng, Bryan Wang, Shao-Hua Sun

专题命中 视频多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.19760 2024-10-29 cs.CV cs.AI cs.MM eess.IV 78%

Movie Trailer Genre Classification Using Multimodal Pretrained Features

Serkan Sulun, Paula Viana, Matthew E. P. Davies

专题命中 视频多模态 :multimodal(title);分类 cs.CV、cs.AI、cs.MM

Journal ref Expert Systems with Applications 258 (2024) 125209

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.08933 2024-10-28 cs.LG physics.ao-ph physics.comp-ph 78%

Multi-Modal Learning-based Reconstruction of High-Resolution Spatial Wind Speed Fields

Matteo Zambra, Nicolas Farrugia, Dorian Cazau, Alexandre Gensse, Ronan Fablet

专题命中 视频多模态 :multi-modal(title,abstract)

Comments 22 pages, 13 figures. This work is to be submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.07238 2024-10-11 cs.HC 78%

vailá: Versatile Anarcho Integrated Liberation Ánalysis in Multimodal Toolbox

Paulo Roberto Pereira Santiago, Abel Gonçalves Chinaglia, Kira Flanagan, Bruno L. S. Bedo, Ligia Yumi Mochida, Juan Aceros, Aline Bononi, Guilherme Manna Cesar

专题命中 视频多模态 :multimodal(title,abstract)

Comments 21 pages, 13 figures, submitted to arXiv under cs.SE (Software Engineering)

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.13394 2024-08-27 cs.RO 78%

Towards Robust Perception for Assistive Robotics: An RGB-Event-LiDAR Dataset and Multi-Modal Detection Pipeline

Adam Scicluna, Cedric Le Gentil, Sheila Sutjipto, Gavin Paul

专题命中 视频多模态 :multi-modal(title,abstract)

Comments Accepted to the 2024 IEEE International Conference on Automation Science and Engineering (CASE)

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.07791 2024-08-16 cs.MM cs.AI cs.CV cs.LG 78%

An Efficient and Explanatory Image and Text Clustering System with Multimodal Autoencoder Architecture

Tiancheng Shi, Yuanchen Wei, John R. Kender

专题命中 视频多模态 :multimodal(title);分类 cs.CV、cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.07488 2024-08-15 cs.HC 78%

Towards Enhanced Context Awareness with Vision-based Multimodal Interfaces

Yongquan Hu, Wen Hu, Aaron Quigley

专题命中 视频多模态 :multimodal(title,abstract)

Comments 3 pages, MOBILEHCI Adjunct '24 26th International Conference on Mobile Human-Computer Interaction, September 30-October 3, 2024, Melbourne, VIC, Australia

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.05255 2024-08-07 cs.HC 78%

Emolysis: A Multimodal Open-Source Group Emotion Analysis and Visualization Toolkit

Shreya Ghosh, Zhixi Cai, Parul Gupta, Garima Sharma, Abhinav Dhall, Munawar Hayat, Tom Gedeon

专题命中 视频多模态 :multimodal(title,abstract)

Comments Accepted by ACII Demo 2024. Both Shreya Ghosh and Zhixi Cai contributed equally to this research

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.17590 2024-06-26 cs.MM cs.AI cs.CV 78%

Multimodal Chaptering for Long-Form TV Newscast Video

Khalil Guetari, Yannis Tevissen, Frédéric Petitpont

专题命中 视频多模态 :multimodal(title);分类 cs.CV、cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏