arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4735 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4735 篇

2507.07261 2025-07-11 cs.LG eess.SP 82%

Robust Multimodal Learning Framework For Intake Gesture Detection Using Contactless Radar and Wearable IMU Sensors

Chunzhuo Wang, Hans Hallez, Bart Vanrumste

机构 * e-Media Research Lab(e-Media研究实验室) ESAT-STADIUS Division, KU Leuven(ESAT-STADIUS部门,KU莱顿大学) M-Group, DistriNet, Department of Computer Science, KU Leuven(M组、DistriNet、计算机科学系,KU莱顿大学)

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract)

Comments This manuscript has been submitted to a peer-reviewed journal and is currently under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07184 2025-06-10 cs.AI cs.CL cs.CV 82%

Mitigating Behavioral Hallucination in Multimodal Large Language Models for Sequential Images

Liangliang You, Junchi Yao, Shu Yang, Guimin Hu, Lijie Hu, Di Wang

机构 * Provable Responsible AI and Data Analytics (PRADA) Lab(可证明负责任的人工智能与数据分析实验室) King Abdullah University of Science and Technology(国王阿卜杜勒阿齐兹大学) University of Electronic Science and Technology of China(中国电子科学技术大学) University of Copenhagen(哥本哈根大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20811 2025-06-10 cs.CV cs.CL cs.MM 82%

HAIC: Improving Human Action Understanding and Generation with Better Captions for Multi-modal Large Language Models

Xiao Wang, Jingyun Hua, Weihong Lin, Yuanxing Zhang, Fuzheng Zhang, Jianlong Wu, Di Zhang, Liqiang Nie

机构 * Kuaishou Technology(快手科技) Shandong University(山东省大学)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.CL、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.02430 2025-06-05 cs.CL cs.AI cs.CV cs.LG 82%

Generative Emotion Cause Explanation in Multimodal Conversations

Lin Wang, Xiaocui Yang, Shi Feng, Daling Wang, Yifei Zhang, Zhitao Zhang

机构 * Northeastern University(东北大学) Shenyang Women’s and Children’s Hospital(沈阳市妇女儿童医院)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01077 2025-06-03 cs.GR cs.HC 82%

TRiMM: Transformer-Based Rich Motion Matching for Real-Time multi-modal Interaction in Digital Humans

Yueqian Guo, Tianzhao Li, Xin Lyu, Jiehaolin Chen, Zhaohan Wang, Sirui Xiao, Yurun Chen, Yezi He, Helin Li, Fan Zhang

专题命中 视频多模态 :multi-modal(title,abstract);cross-modal(abstract)

Comments 24 pages,12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.01932 2025-05-29 cs.RO cs.LG 82%

Bridging Language, Vision and Action: Multimodal VAEs in Robotic Manipulation Tasks

Gabriela Sejnova, Michal Vavrecka, Karla Stepanova

机构 * Czech Institute of Informatics, Robotics and Cybernetics(捷克信息学、机器人学与自动控制研究所) Czech Technical University in Prague(布拉格捷克技术大学)

专题命中 视频多模态 :multimodal(title,abstract);image-text(abstract)

Comments 7 pages, 5 figures, 2 tables, conference

Journal ref 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14535 2025-05-21 cs.LG cs.HC 82%

Spiking Neural Networks with Temporal Attention-Guided Adaptive Fusion for imbalanced Multi-modal Learning

Jiangrong Shen, Yulin Xie, Qi Xu, Gang Pan, Huajin Tang, Badong Chen

机构 * Faculty of Electronic and Information Engineering, Xi’an Jiaotong University(电子与信息工程学院,西安交通大学) Institute of Artificial Intelligence and Robotics, Xi’an Jiaotong University(人工智能与机器人研究院,西安交通大学) State Key Lab of Brain-Machine Intelligence, Zhejiang University(脑机智能国家重点实验室,浙江大学) School of Computer Science, Dalian University of Technology(计算机科学学院,大连理工大学) National Key Lab of Human-Machine Hybrid Augmented Intelligence, Xi’an Jiaotong University(人机混合增强智能国家实验室,西安交通大学)

专题命中 视频多模态 :multi-modal(title);multimodal(abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12051 2025-05-20 cs.MM cs.AI cs.CV 82%

Enhanced Multimodal Hate Video Detection via Channel-wise and Modality-wise Fusion

Yinghui Zhang, Tailin Chen, Yuchen Zhang, Zeyu Fu

机构 * Department of Computer Science, University of Exeter(埃克塞特大学计算机科学系) Institute for Analytics and Data Science, University of Essex(埃塞克斯大学分析与数据科学研究所)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI、cs.MM

Comments ICDMW 2024, Github: https://github.com/EvelynZ10/cmfusion

Journal ref 2024 IEEE International Conference on Data Mining Workshops (ICDMW), Abu Dhabi, United Arab Emirates, 2024, pp. 183-190

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16036 2025-03-21 cs.CV cs.AI cs.CL 82%

Hybrid-Level Instruction Injection for Video Token Compression in Multi-modal Large Language Models

Zhihang Liu, Chen-Wei Xie, Pandeng Li, Liming Zhao, Longxiang Tang, Yun Zheng, Chuanbin Liu, Hongtao Xie

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted to CVPR2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.05889 2025-03-21 cs.CV cs.AI cs.CL 82%

CREMA: Generalizable and Efficient Video-Language Reasoning via Multimodal Modular Fusion

Shoubin Yu, Jaehong Yoon, Mohit Bansal

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments ICLR 2025; first two authors contributed equally. Project page: https://CREMA-VideoLLM.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.05092 2025-03-19 cs.CV cs.AI cs.CL 82%

Lost in Time: Clock and Calendar Understanding Challenges in Multimodal LLMs

Rohit Saxena, Aryo Pradipta Gema, Pasquale Minervini

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted at the ICLR 2025 Workshop on Reasoning and Planning for Large Language Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.16755 2025-02-25 cs.CY 82%

Watch Out E-scooter Coming Through: Multimodal Sensing of Mixed Traffic Use and Conflicts Through Riders Ego-centric Views

Hiruni Nuwanthika Kegalle, Danula Hettiachchi, Jeffrey Chan, Mark Sanderson, Flora D. Salim

专题命中 视频多模态 :multimodal(title,abstract);multi-modal(abstract)

Comments Accepted in Proc. ACM Interactive, Mobile, Wearable and Ubiquitous Technologies,(March 2025), 23 pages. https://doi.org/10.1145/3712284

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.05474 2025-01-13 cs.CL cs.AI cs.LG cs.SD eess.AS 82%

Modality-Invariant Bidirectional Temporal Representation Distillation Network for Missing Multimodal Sentiment Analysis

Xincheng Wang, Liejun Wang, Yinfeng Yu, Xinxin Jiao

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI、eess.AS

Comments Accepted for publication by 2025 IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.18748 2025-01-03 cs.MM cs.CL cs.SD eess.AS 82%

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction

Yuan Zhao, Rui Liu, Gaoxiang Cong

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL、cs.MM、eess.AS

Comments Accepted by ICASSP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.16407 2024-10-23 cs.CL cs.AI cs.MM 82%

Enhancing Multimodal Affective Analysis with Learned Live Comment Features

Zhaoyuan Deng, Amith Ananthram, Kathleen McKeown

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.10818 2024-10-16 cs.CV cs.AI cs.CL cs.LG 82%

TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Mu Cai, Reuben Tan, Jianrui Zhang, Bocheng Zou, Kai Zhang, Feng Yao, Fangrui Zhu, Jing Gu, Yiwu Zhong, Yuzhang Shang, Yao Dou, Jaden Park, Jianfeng Gao, Yong Jae Lee, Jianwei Yang

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Project Page: https://temporalbench.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.19467 2024-10-11 cs.CL cs.AI cs.CV 82%

TV-TREES: Multimodal Entailment Trees for Neuro-Symbolic Video Reasoning

Kate Sanders, Nathaniel Weir, Benjamin Van Durme

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments 9 pages, EMNLP 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.15766 2024-10-04 cs.AI cs.CL cs.CV 82%

Enhancing Adverse Drug Event Detection with Multimodal Dataset: Corpus Creation and Model Development

Pranab Sahoo, Ayush Kumar Singh, Sriparna Saha, Aman Chadha, Samrat Mondal

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments ACL Findings 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.11593 2024-09-05 cs.MM cs.CV cs.SD eess.AS 82%

MCDubber: Multimodal Context-Aware Expressive Video Dubbing

Yuan Zhao, Zhenqi Jia, Rui Liu, De Hu, Feilong Bao, Guanglai Gao

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM、eess.AS

Comments Accepted by NCMMSC2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.14895 2024-08-29 cs.AI cs.CL cs.CV 82%

VHAKG: A Multi-modal Knowledge Graph Based on Synchronized Multi-view Videos of Daily Activities

Shusaku Egami, Takahiro Ugai, Swe Nwe Nwe Htun, Ken Fukuda

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments 5 pages, 4 figures, accepted by CIKM2024 Resource Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.07694 2024-08-15 cs.CV cs.AI cs.LG cs.MM 82%

End-to-end Semantic-centric Video-based Multimodal Affective Computing

Ronghao Lin, Ying Zeng, Sijie Mai, Haifeng Hu

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI、cs.MM

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.10711 2024-07-24 cs.CV cs.AI cs.CL 82%

Weakly Supervised Gaussian Contrastive Grounding with Large Multimodal Models for Video Question Answering

Haibo Wang, Chenghang Lai, Yixuan Sun, Weifeng Ge

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments accepted by ACM Multimedia 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.08743 2024-06-18 cs.AI cs.CL cs.CV cs.LG 82%

MMToM-QA: Multimodal Theory of Mind Question Answering

Chuanyang Jin, Yutong Wu, Jing Cao, Jiannan Xiang, Yen-Ling Kuo, Zhiting Hu, Tomer Ullman, Antonio Torralba, Joshua B. Tenenbaum, Tianmin Shu

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments ACL 2024. 26 pages, 11 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.06964 2024-06-12 cs.CL cs.MM cs.SD eess.AS 82%

Missingness-resilient Video-enhanced Multimodal Disfluency Detection

Payal Mohapatra, Shamika Likhite, Subrata Biswas, Bashima Islam, Qi Zhu

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL、cs.MM、eess.AS

Comments Accepted to Interspeech 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.16728 2024-05-28 cs.CV cs.AI cs.LG cs.MM 82%

Towards Multi-Task Multi-Modal Models: A Video Generative Perspective

Lijun Yu

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.AI、cs.MM

Comments PhD thesis

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.05291 2024-04-01 cs.CV cs.AI cs.CL 82%

GlitchBench: Can large multimodal models detect video game glitches?

Mohammad Reza Taesiri, Tianjun Feng, Anh Nguyen, Cor-Paul Bezemer

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments CVPR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.02051 2024-03-29 cs.CV cs.AI cs.CL 82%

TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Shuhuai Ren, Linli Yao, Shicheng Li, Xu Sun, Lu Hou

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments CVPR 2024 camera-ready version, code is available at https://github.com/RenShuhuai-Andy/TimeChat

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.17172 2023-12-29 cs.CV cs.AI cs.CL 82%

Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

Jiasen Lu, Christopher Clark, Sangho Lee, Zichen Zhang, Savya Khosla, Ryan Marten, Derek Hoiem, Aniruddha Kembhavi

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments 38 pages, 20 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.00347 2023-12-19 cs.CV cs.CL cs.MM 82%

RTQ: Rethinking Video-language Understanding Based on Image-text Model

Xiao Wang, Yaoyu Li, Tian Gan, Zheng Zhang, Jingjing Lv, Liqiang Nie

专题命中 视频多模态 :image-text(title,abstract);分类 cs.CV、cs.CL、cs.MM

Comments Accepted by ACM MM 2023 as Oral representation

Journal ref In International Conference on Multimedia. ACM, 557--566 (2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.03741 2023-08-08 cs.CV cs.AI cs.LG cs.MM 82%

MAiVAR-T: Multimodal Audio-image and Video Action Recognizer using Transformers

Muhammad Bilal Shaikh, Douglas Chai, Syed Mohammed Shamsul Islam, Naveed Akhtar

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI、cs.MM

Comments 6 pages, 7 figures, 4 tables, Peer reviewed, Accepted @ The 11th European Workshop on Visual Information Processing (EUVIP) will be held on 11th-14th September 2023, in Gjøvik, Norway. arXiv admin note: text overlap with arXiv:2103.15691 by other authors

详情

展开后加载摘要…

URL PDF HTML 收藏