arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4757 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4757 篇

2503.07963 2025-10-06 cs.RO cs.AI cs.SY eess.SY 74%

Hierarchical Contact-Rich Trajectory Optimization for Multi-Modal Manipulation using Tight Convex Relaxations

Yuki Shirai, Arvind Raghunathan, Devesh K. Jha

机构 * Mitsubishi Electric Research Laboratories(三菱电机研究实验室)

专题命中 视频多模态 :multi-modal(title);分类 cs.AI

Comments 2025 IEEE International Conference on Robotics and Automation (2025 ICRA)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.00139 2025-10-01 cs.CV 74%

SuperEvent: Cross-Modal Learning of Event-based Keypoint Detection for SLAM

Yannick Burkhardt, Simon Schaefer, Stefan Leutenegger

机构 * Technical University of Munich(慕尼黑技术大学) ETH Zürich(苏黎世联邦理工学院) Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心)

专题命中 视频多模态 :cross-modal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.13838 2025-09-30 eess.SP cs.CV cs.IT eess.IV math.IT 74%

Generative Video Semantic Communication via Multimodal Semantic Fusion with Large Model

Hang Yin, Li Qiao, Yu Ma, Shuo Sun, Kan Li, Zhen Gao, Dusit Niyato

机构 * School of Information and Electronics, Beijing Institute of Technology(信息与电子学院,北京理工大学) School of Computer Science and Engineering, Nanyang Technological University(计算机科学与工程学院,南洋理工大学)

专题命中 视频多模态 :multimodal(title);分类 cs.CV

Comments IEEE Transactions on Vehicular Technology

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.12341 2025-08-14 cs.MM 74%

Multimodal LLM-based Query Paraphrasing for Video Search

Jiaxin Wu, Chong-Wah Ngo, Wing-Kwong Chan, Sheng-Hua Zhong, Xiong-Yong Wei, Qing Li

专题命中 视频多模态 :multimodal(title);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23524 2025-06-06 cs.CV 74%

CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization

Rui Xia, Dan Jiang, Quan Zhang, Ke Zhang, Chun Yuan

专题命中 视频多模态 :audio-visual(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.02252 2025-06-03 cs.CV 74%

MoviePuzzle: Visual Narrative Reasoning through Multimodal Order Learning

Jianghui Wang, Yuxuan Wang, Dongyan Zhao, Zilong Zheng

机构 * Beijing Institute for General Artificial Intelligence(北京通用人工智能研究院) Wangxuan Institute of Computer Technology(王轩计算机技术研究所)

专题命中 视频多模态 :multimodal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05541 2025-04-10 cs.CV 74%

Caption Anything in Video: Fine-grained Object-centric Captioning via Spatiotemporal Multimodal Prompting

Yunlong Tang, Jing Bi, Chao Huang, Susan Liang, Daiki Shimada, Hang Hua, Yunzhong Xiao, Yizhi Song, Pinxin Liu, Mingqian Feng, Junjia Guo, Zhuo Liu, Luchuan Song, Ali Vosoughi, Jinxi He, Liu He, Zeliang Zhang, Jiebo Luo, Chenliang Xu

专题命中 视频多模态 :multimodal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05673 2025-04-09 cs.CV 74%

VC-LLM: Automated Advertisement Video Creation from Raw Footage using Multi-modal LLMs

Dongjun Qian, Kai Su, Yiming Tan, Qishuai Diao, Xian Wu, Chang Liu, Bingyue Peng, Zehuan Yuan

专题命中 视频多模态 :multi-modal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.03112 2025-03-06 cs.SI cs.AI cs.NE 74%

A Multimodal Framework for Topic Propagation Classification in Social Networks

Yuchuan Jiang, Chaolong Jia, Yunyi Qin, Wei Cai, Yongsen Qian

专题命中 视频多模态 :multimodal(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.17805 2024-12-24 cs.CV 74%

Large Motion Video Autoencoding with Cross-modal Video VAE

Yazhou Xing, Yang Fei, Yingqing He, Jingye Chen, Jiaxin Xie, Xiaowei Chi, Qifeng Chen

专题命中 视频多模态 :cross-modal(title);分类 cs.CV

Comments Project Website: https://yzxing87.github.io/vae/

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.00517 2024-12-03 cs.AI cs.ET cs.RO 74%

LAMBDA: Covering the Multimodal Critical Scenarios for Automated Driving Systems by Search Space Quantization

Xinzheng Wu, Junyi Chen, Xingyu Xing, Jian Sun, Ye Tian, Lihao Liu, Yong Shen

专题命中 视频多模态 :multimodal(title);分类 cs.AI

Comments 17pages, 21figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.19825 2024-10-29 cs.CV 74%

Automating Video Thumbnails Selection and Generation with Multimodal and Multistage Analysis

Elia Fantini

专题命中 视频多模态 :multimodal(title);分类 cs.CV

Comments 150 pages, 60 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.16592 2024-10-23 cs.LG cs.CL cs.CY 74%

ViMGuard: A Novel Multi-Modal System for Video Misinformation Guarding

Andrew Kan, Christopher Kan, Zaid Nabulsi

专题命中 视频多模态 :multi-modal(title);分类 cs.CL

Comments 7 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.12420 2024-10-03 cs.CL cs.LG 74%

MMUTF: Multimodal Multimedia Event Argument Extraction with Unified Template Filling

Philipp Seeberger, Dominik Wagner, Korbinian Riedhammer

专题命中 视频多模态 :multimodal(title);分类 cs.CL

Comments Accepted to Findings of EMNLP 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.00017 2024-10-02 cs.CV eess.SP 74%

Multimodal Power Outage Prediction for Rapid Disaster Response and Resource Allocation

Alejandro Aparcedo, Christian Lopez, Abhinav Kotta, Mengjie Li

专题命中 视频多模态 :multimodal(title);分类 cs.CV

Comments 7 pages, 4 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.02690 2024-09-05 cs.SI cs.CL 74%

Detecting Calls to Action in Multimodal Content: Analysis of the 2021 German Federal Election Campaign on Instagram

Michael Achmann-Denkler, Jakob Fehle, Mario Haim, Christian Wolff

专题命中 视频多模态 :multimodal(title);分类 cs.CL

Comments Accepted Archival Paper for the CPSS Workshop at KONVENS 2024. Camera Ready Submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.14930 2024-08-29 cs.CV 74%

CMTA: Cross-Modal Temporal Alignment for Event-guided Video Deblurring

Taewoo Kim, Hoonhee Cho, Kuk-Jin Yoon

专题命中 视频多模态 :cross-modal(title);分类 cs.CV

Comments Accepted in ECCV2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.15377 2024-08-15 cs.CV 74%

InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Yi Wang, Kunchang Li, Xinhao Li, Jiashuo Yu, Yinan He, Chenting Wang, Guo Chen, Baoqi Pei, Ziang Yan, Rongkun Zheng, Jilan Xu, Zun Wang, Yansong Shi, Tianxiang Jiang, Songze Li, Hongjie Zhang, Yifei Huang, Yu Qiao, Yali Wang, Limin Wang

专题命中 视频多模态 :multimodal(title);分类 cs.CV

Comments a technical report about video understanding (accepted to ECCV2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.15921 2024-07-02 cs.CV 74%

PUDD: Towards Robust Multi-modal Prototype-based Deepfake Detection

Alvaro Lopez Pellcier, Yi Li, Plamen Angelov

专题命中 视频多模态 :multi-modal(title);分类 cs.CV

Comments CVPR2024

Journal ref CVPR2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.09585 2024-06-12 cs.CV 74%

StreamingFlow: Streaming Occupancy Forecasting with Asynchronous Multi-modal Data Streams via Neural Ordinary Differential Equation

Yining Shi, Kun Jiang, Ke Wang, Jiusi Li, Yunlong Wang, Mengmeng Yang, Diange Yang

专题命中 视频多模态 :multi-modal(title);分类 cs.CV

Comments cvpr2024 poster (highlight), code at https://github.com/synsin0/StreamingFlow

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.00822 2024-06-12 cs.AI cs.HC cs.RO 74%

Open-Ended Multi-Modal Relational Reasoning for Video Question Answering

Haozheng Luo, Ruiyang Qin, Chenwei Xu, Guo Ye, Zening Luo

专题命中 视频多模态 :multi-modal(title);分类 cs.AI

Comments 2023 32nd IEEE International Conference on Robot and Human Interactive Communication (RO-MAN)

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.10825 2024-03-19 cs.CV 74%

Affective Behaviour Analysis via Integrating Multi-Modal Knowledge

Wei Zhang, Feng Qiu, Chen Liu, Lincheng Li, Heming Du, Tiancheng Guo, Xin Yu

专题命中 视频多模态 :multi-modal(title);分类 cs.CV

Comments 11 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.08369 2024-02-14 cs.AI 74%

One-shot Imitation in a Non-Stationary Environment via Multi-Modal Skill

Sangwoo Shin, Daehee Lee, Minjong Yoo, Woo Kyung Kim, Honguk Woo

专题命中 视频多模态 :multi-modal(title);分类 cs.AI

Comments ICML-2023 Camera Ready Version

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.12419 2024-01-24 cs.CV 74%

Multi-modal News Understanding with Professionally Labelled Videos (ReutersViLNews)

Shih-Han Chou, Matthew Kowal, Yasmin Niknam, Diana Moyano, Shayaan Mehdi, Richard Pito, Cheng Zhang, Ian Knopke, Sedef Akinli Kocak, Leonid Sigal, Yalda Mohsenzadeh

专题命中 视频多模态 :multi-modal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.14634 2023-12-25 cs.RO cs.AI 74%

Mining multi-modal communication patterns in interaction with explainable and non-explainable robots

Suna Bensch, Amanda Eriksson

专题命中 视频多模态 :multi-modal(title);分类 cs.AI

Journal ref IEEE RO-MAN 2023, 32nd IEEE International conference on Robot and Human Interactive Communication; Workshop Human-Robot Interaction for Explainability in Robotics, Busan, Korea, August 28-31, 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.01910 2023-11-27 cs.CV 74%

Multimodal Generation of Novel Action Appearances for Synthetic-to-Real Recognition of Activities of Daily Living

Zdravko Marinov, David Schneider, Alina Roitberg, Rainer Stiefelhagen

专题命中 视频多模态 :multimodal(title);分类 cs.CV

Comments 8 pages, 7 figures, to be published in IROS 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.05494 2023-11-10 cs.CV cs.RO 74%

Object-centric Cross-modal Feature Distillation for Event-based Object Detection

Lei Li, Alexander Liniger, Mario Millhaeusler, Vagia Tsiminaki, Yuanyou Li, Dengxin Dai

专题命中 视频多模态 :cross-modal(title);分类 cs.CV

Comments 12 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1712.00796 2023-09-26 cs.CV 74%

Multimodal Visual Concept Learning with Weakly Supervised Techniques

Giorgos Bouritsas, Petros Koutras, Athanasia Zlatintsi, Petros Maragos

专题命中 视频多模态 :multimodal(title);分类 cs.CV

Comments CVPR 2018

Journal ref Proc. IEEE/CVF Conf. Comp. Vis. Patt. Rec. (CVPR) pp. 4914 - 4923 (2018)

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.09306 2023-07-19 cs.CV cs.LG cs.RO 74%

EigenTrajectory: Low-Rank Descriptors for Multi-Modal Trajectory Forecasting

Inhwan Bae, Jean Oh, Hae-Gon Jeon

专题命中 视频多模态 :multi-modal(title);分类 cs.CV

Comments Accepted at ICCV 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.08063 2023-03-21 cs.CV 74%

MINOTAUR: Multi-task Video Grounding From Multimodal Queries

Raghav Goyal, Effrosyni Mavroudi, Xitong Yang, Sainbayar Sukhbaatar, Leonid Sigal, Matt Feiszli, Lorenzo Torresani, Du Tran

专题命中 视频多模态 :multimodal(title);分类 cs.CV

Comments 22 pages, 8 figures and 13 tables

详情

展开后加载摘要…

URL PDF HTML 收藏