arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4749 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4749 篇

2411.04923 2025-03-26 cs.CV 79%

VideoGLaMM: A Large Multimodal Model for Pixel-Level Visual Grounding in Videos

Shehan Munasinghe, Hanan Gani, Wenqi Zhu, Jiale Cao, Eric Xing, Fahad Shahbaz Khan, Salman Khan

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments Technical Report of VideoGLaMM

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18933 2025-03-25 cs.CV 79%

SyncVP: Joint Diffusion for Synchronous Multi-Modal Video Prediction

Enrico Pallotta, Sina Mokhtarzadeh Azar, Shuai Li, Olga Zatsarynna, Juergen Gall

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.17827 2025-03-25 cs.CV 79%

4D-Bench: Benchmarking Multi-modal Large Language Models for 4D Object Understanding

Wenxuan Zhu, Bing Li, Cheng Zheng, Jinjie Mai, Jun Chen, Letian Jiang, Abdullah Hamdi, Sara Rojas Martinez, Chia-Wen Lin, Mohamed Elhoseiny, Bernard Ghanem

专题命中 视频多模态 :multi-modal(title);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16466 2025-03-24 cs.HC cs.AI 79%

ACE, Action and Control via Explanations: A Proposal for LLMs to Provide Human-Centered Explainability for Multimodal AI Assistants

Elizabeth Anne Watkins, Emanuel Moss, Ramesh Manuvinakurike, Meng Shi, Richard Beckwith, Giuseppe Raffa

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

Comments Accepted at Human-Centered Explainable AI workshop at CHI 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.15647 2025-03-21 cs.CV cs.LG 79%

Multi-Modal Gesture Recognition from Video and Surgical Tool Pose Information via Motion Invariants

Jumanh Atoum, Garrison L. H. Johnston, Nabil Simaan, Jie Ying Wu

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.14498 2025-03-19 cs.CV cs.RO 79%

Tracking Meets Large Multimodal Models for Driving Scenario Understanding

Ayesha Ishaq, Jean Lahoud, Fahad Shahbaz Khan, Salman Khan, Hisham Cholakkal, Rao Muhammad Anwer

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments 13 pages, 8 figures, Github: https://github.com/mbzuai-oryx/TrackingMeetsLMM

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.13646 2025-03-19 cs.CV 79%

Omnia de EgoTempo: Benchmarking Temporal Understanding of Multi-Modal LLMs in Egocentric Videos

Chiara Plizzari, Alessio Tonioni, Yongqin Xian, Achin Kulshrestha, Federico Tombari

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted to CVPR 2025. Dataset and code are available at https://github.com/google-research-datasets/egotempo.git

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.13016 2025-03-18 cs.CV 79%

Efficient Motion-Aware Video MLLM

Zijia Zhao, Yuqi Huo, Tongtian Yue, Longteng Guo, Haoyu Lu, Bingning Wang, Weipeng Chen, Jing Liu

专题命中 视频多模态 :MLLM(title,abstract);分类 cs.CV

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.11695 2025-03-18 cs.LG cs.AI 79%

MELON: Multimodal Mixture-of-Experts with Spectral-Temporal Fusion for Long-Term Mobility Estimation in Critical Care

Jiaqing Zhang, Miguel Contreras, Jessica Sena, Andrea Davidson, Yuanfang Ren, Ziyuan Guan, Tezcan Ozrazgat-Baslanti, Tyler J. Loftus, Subhash Nerella, Azra Bihorac, Parisa Rashidi

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.02063 2025-03-17 cs.CV 79%

V$^2$Dial: Unification of Video and Visual Dialog via Multimodal Experts

Adnen Abdessaied, Anna Rohrbach, Marcus Rohrbach, Andreas Bulling

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.15363 2025-03-17 cs.HC cs.CV 79%

M2LADS Demo: A System for Generating Multimodal Learning Analytics Dashboards

Alvaro Becerra, Roberto Daza, Ruth Cobos, Aythami Morales, Julian Fierrez

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments Published in the Workshop on Innovation and Responsibility in AI-Supported Education (iRAISE25) at AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.08370 2025-03-12 cs.GR cs.CV 79%

Ev-Layout: A Large-scale Event-based Multi-modal Dataset for Indoor Layout Estimation and Tracking

Xucheng Guo, Yiran Shen, Xiaofang Xiao, Yuanfeng Zhou, Lin Wang

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.05936 2025-03-11 cs.CV 79%

CASP: Compression of Large Multimodal Models Based on Attention Sparsity

Mohsen Gholami, Mohammad Akbari, Kevin Cannons, Yong Zhang

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.15393 2025-03-04 cs.CV 79%

LongCaptioning: Unlocking the Power of Long Video Caption Generation in Large Multimodal Models

Hongchen Wei, Zhihong Tan, Yaosi Hu, Chang Wen Chen, Zhenzhong Chen

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.18371 2025-02-26 cs.AI 79%

MindMem: Multimodal for Predicting Advertisement Memorability Using LLMs and Deep Learning

Sepehr Asgarian, Qayam Jetha, Jouhyun Jeon

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

Comments 7 pages, 5 figures, 4 Tables, AAAI 2025 Economics of Modern ML: Markets, Incentives, and Generative AI Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.17038 2025-02-25 cs.MM 79%

Multi-modal and Metadata Capture Model for Micro Video Popularity Prediction

Jiacheng Lu, Mingyuan Xiao, Weijian Wang, Yuxin Du, Zhengze Wu, Cheng Hua

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.17359 2025-02-24 cs.CL 79%

Chain-of-Action: Faithful and Multimodal Question Answering through Large Language Models

Zhenyu Pan, Haozheng Luo, Manling Li, Han Liu

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL

Comments International Conference on Learning Representations (ICLR) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.14227 2025-02-21 cs.LG cs.AI 79%

SleepGMUformer: A gated multimodal temporal neural network for sleep staging

Chenjun Zhao, Xuesen Niu, Xinglin Yu, Long Chen, Na Lv, Huiyu Zhou, Aite Zhao

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.13716 2025-02-20 cs.CV 79%

Event-Based Video Frame Interpolation With Cross-Modal Asymmetric Bidirectional Motion Fields

Taewoo Kim, Yujeong Chae, Hyun-Kurl Jang, Kuk-Jin Yoon

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

Comments Accepted in CVPR2023(Highlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.12604 2025-02-19 cs.CV 79%

S2C: Learning Noise-Resistant Differences for Unsupervised Change Detection in Multimodal Remote Sensing Images

Lei Ding, Xibing Zuo, Danfeng Hong, Haitao Guo, Jun Lu, Zhihui Gong, Lorenzo Bruzzone

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.06332 2025-02-18 cs.CV 79%

X-VARS: Introducing Explainability in Football Refereeing with Multi-Modal Large Language Model

Jan Held, Hani Itani, Anthony Cioppa, Silvio Giancola, Bernard Ghanem, Marc Van Droogenbroeck

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.09895 2025-02-11 cs.CV 79%

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP

Yating Yu, Congqi Cao, Yueran Zhang, Qinyi Lv, Lingtong Min, Yanning Zhang

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted by AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.14269 2025-01-31 cs.IR cs.AI 79%

Hierarchical Time-Aware Mixture of Experts for Multi-Modal Sequential Recommendation

Shengzhe Zhang, Liyi Chen, Dazhong Shen, Chao Wang, Hui Xiong

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.AI

Comments Accepted to WWW 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.15508 2025-01-28 cs.MM 79%

Learning Complex Heterogeneous Multimodal Fake News via Social Latent Network Inference

Mingxin Li, Yuchen Zhang, Haowei Xu, Xianghua Li, Chao Gao, Zhen Wang

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.MM

Comments AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.10692 2025-01-22 cs.CV 79%

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection

Yifang Xu, Yunzhuo Sun, Benxiang Zhai, Zien Xie, Youyao Jia, Sidan Du

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted by ICME 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.09499 2025-01-17 cs.CV 79%

VanGogh: A Unified Multimodal Diffusion-based Framework for Video Colorization

Zixun Fang, Zhiheng Liu, Kai Zhu, Yu Liu, Ka Leong Cheng, Wei Zhai, Yang Cao, Zheng-Jun Zha

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.06475 2025-01-14 cs.CV cs.LG 79%

Enhancing Multi-Modal Video Sentiment Classification Through Semi-Supervised Clustering

Mehrshad Saadatinia, Minoo Ahmadi, Armin Abdollahi

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.05884 2025-01-13 cs.CV 79%

Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs

Dabing Cheng, Haosen Zhan, Xingchen Zhao, Guisheng Liu, Zemin Li, Jinghui Xie, Zhao Song, Weiguo Feng, Bingyue Peng

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments 16pages conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.05733 2025-01-13 cs.CV 79%

TB-Bench: Training and Testing Multi-Modal AI for Understanding Spatio-Temporal Traffic Behaviors from Dashcam Images/Videos

Korawat Charoenpitaks, Van-Quang Nguyen, Masanori Suganuma, Kentaro Arai, Seiji Totsuka, Hiroshi Ino, Takayuki Okatani

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments Main Paper: 8 pages, Supplementary Materials: 15 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.11280 2025-01-09 cs.CV 79%

ISR-DPO: Aligning Large Multimodal Models for Videos by Iterative Self-Retrospective DPO

Daechul Ahn, Yura Choi, San Kim, Youngjae Yu, Dongyeop Kang, Jonghyun Choi

专题命中 视频多模态 :multimodal(title);multi-modal(abstract);分类 cs.CV

Comments AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏