arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4749 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4749 篇

2305.12561 2023-05-23 cs.HC cs.CV 79%

M2LADS: A System for Generating MultiModal Learning Analytics Dashboards in Open Education

Álvaro Becerra, Roberto Daza, Ruth Cobos, Aythami Morales, Mutlu Cukurova, Julian Fierrez

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments Accepted in "Workshop on Open Education Resources (OER) of COMPSAC 2023"

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.08401 2023-05-18 cs.CV cs.LG 79%

Multimodal Short Video Rumor Detection System Based on Contrastive Learning

Yuxing Yang, Junhao Zhao, Siyi Wang, Xiangyu Min, Pengchao Wang, Haizhou Wang

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.14407 2023-05-02 cs.CV 79%

ChatVideo: A Tracklet-centric Multimodal and Versatile Video Understanding System

Junke Wang, Dongdong Chen, Chong Luo, Xiyang Dai, Lu Yuan, Zuxuan Wu, Yu-Gang Jiang

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.10590 2023-04-19 cs.CV 79%

Multi-modal Facial Action Unit Detection with Large Pre-trained Models for the 5th Competition on Affective Behavior Analysis in-the-wild

Yufeng Yin, Minh Tran, Di Chang, Xinrui Wang, Mohammad Soleymani

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments 8 pages, 7 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.10849 2023-04-12 cs.CV 79%

Multi-modal Facial Affective Analysis based on Masked Autoencoder

Wei Zhang, Bowen Ma, Feng Qiu, Yu Ding

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.11732 2023-03-22 cs.CV 79%

Multi-modal Prompting for Low-Shot Temporal Action Localization

Chen Ju, Zeqian Li, Peisen Zhao, Ya Zhang, Xiaopeng Zhang, Qi Tian, Yanfeng Wang, Weidi Xie

专题命中 视频多模态 :multi-modal(title);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.02136 2023-03-07 cs.CV 79%

Efficient End-to-End Video Question Answering with Pyramidal Multimodal Transformer

Min Peng, Chongyang Wang, Yu Shi, Xiang-Dong Zhou

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments Accepted by AAAI 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.06031 2023-03-03 cs.CV 79%

Long-Form Video-Language Pre-Training with Multimodal Temporal Contrastive Learning

Yuchong Sun, Hongwei Xue, Ruihua Song, Bei Liu, Huan Yang, Jianlong Fu

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments Accepted by NeurIPS 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.10465 2023-02-22 cs.CV 79%

A Flexible Multi-view Multi-modal Imaging System for Outdoor Scenes

Meng Zhang, Wenxuan Guo, Bohao Fan, Yifan Chen, Jianjiang Feng, Jie Zhou

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.11314 2023-02-22 cs.CV 79%

Modality Mixer for Multi-modal Action Recognition

Sumin Lee, Sangmin Woo, Yeonju Park, Muhammad Adi Nugroho, Changick Kim

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.00912 2023-02-08 cs.CV 79%

Advances and Challenges in Multimodal Remote Sensing Image Registration

Bai Zhu, Liang Zhou, Simiao Pu, Jianwei Fan, Yuanxin Ye

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments 10 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.04780 2023-01-20 cs.CV 79%

MAiVAR: Multimodal Audio-Image and Video Action Recognizer

Muhammad Bilal Shaikh, Douglas Chai, Syed Mohammed Shamsul Islam, Naveed Akhtar

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments Peer reviewed & accepted at IEEE VCIP 2022 (http://www.vcip2022.org/)

Journal ref 2022 IEEE International Conference on Visual Communications and Image Processing (VCIP)

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.00167 2023-01-13 cs.CV 79%

Event-Based Fusion for Motion Deblurring with Cross-modal Attention

Lei Sun, Christos Sakaridis, Jingyun Liang, Qi Jiang, Kailun Yang, Peng Sun, Yaozu Ye, Kaiwei Wang, Luc Van Gool

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

Comments Accepted by ECCV 2022 as oral presentation

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.10596 2023-01-12 cs.CV 79%

Open-Vocabulary Temporal Action Detection with Off-the-Shelf Image-Text Features

Vivek Rathod, Bryan Seybold, Sudheendra Vijayanarasimhan, Austin Myers, Xiuye Gu, Vighnesh Birodkar, David A. Ross

专题命中 视频多模态 :image-text(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.02177 2023-01-05 cs.LG cs.CL 79%

GCNet: Graph Completion Network for Incomplete Multimodal Learning in Conversation

Zheng Lian, Lan Chen, Licai Sun, Bin Liu, Jianhua Tao

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.14143 2023-01-02 cs.CV 79%

Multimodal Wildland Fire Smoke Detection

Siddhant Baldota, Shreyas Anantha Ramaprasad, Jaspreet Kaur Bhamra, Shane Luna, Ravi Ramachandra, Eugene Zen, Harrison Kim, Daniel Crawl, Ismael Perez, Ilkay Altintas, Garrison W. Cottrell, Mai H. Nguyen

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.06345 2022-12-27 cs.MM 79%

Self-supervised Multi-Modal Video Forgery Attack Detection

Chenhui Zhao, Xiang Li, Rabih Younes

专题命中 视频多模态 :multi-modal(title);multimodal(abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.09522 2022-12-20 cs.CV 79%

MIST: Multi-modal Iterative Spatial-Temporal Transformer for Long-form Video Question Answering

Difei Gao, Luowei Zhou, Lei Ji, Linchao Zhu, Yi Yang, Mike Zheng Shou

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.08859 2022-12-20 cs.RO cs.CV cs.LG 79%

iCub! Do you recognize what I am doing?: multimodal human action recognition on multisensory-enabled iCub robot

Kas Kniesmeijer, Murat Kirtay

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments 7 pages, 5 figures and 1 table. International Conference on Social Robotics

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.04700 2022-12-12 cs.CV 79%

Tencent AVS: A Holistic Ads Video Dataset for Multi-modal Scene Segmentation

Jie Jiang, Zhimin Li, Jiangfeng Xiong, Rongwei Quan, Qinglin Lu, Wei Liu

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.13804 2022-11-28 cs.HC cs.CL 79%

On the Linguistic and Computational Requirements for Creating Face-to-Face Multimodal Human-Machine Interaction

João Ranhel, Cacilda Vilela de Lima

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.03785 2022-11-08 cs.AI cs.RO 79%

Learning Visual Locomotion with Cross-Modal Supervision

Antonio Loquercio, Ashish Kumar, Jitendra Malik

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.AI

Comments Learning to walk from pixels in the real world by using proprioception as supervision. Project page for videos and code: https://antonilo.github.io/vision_locomotion/

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.06125 2022-11-04 cs.LG cs.AI physics.ao-ph 79%

Hurricane Forecasting: A Novel Multimodal Machine Learning Framework

Léonard Boussioux, Cynthia Zeng, Théo Guénais, Dimitris Bertsimas

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

Comments Published by the AMS' Weather and Forecasting journal; Spotlight talk at NeurIPS 2021, Tackling Climate Change with AI ; https://journals.ametsoc.org/view/journals/wefo/37/6/WAF-D-21-0091.1.xml

Journal ref 2022, Weather and Forecasting, 37(6), 817-831

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.14512 2022-10-27 cs.CV 79%

End-to-End Multimodal Representation Learning for Video Dialog

Huda Alamri, Anthony Bilic, Michael Hu, Apoorva Beedu, Irfan Essa

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.12649 2022-10-25 cs.CV cs.RO 79%

Anticipative Feature Fusion Transformer for Multi-Modal Action Anticipation

Zeyun Zhong, David Schneider, Michael Voit, Rainer Stiefelhagen, Jürgen Beyerer

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted to WACV 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.08452 2022-10-18 cs.CV 79%

Efficient Cross-Modal Video Retrieval with Meta-Optimized Frames

Ning Han, Xun Yang, Ee-Peng Lim, Hao Chen, Qianru Sun

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.07886 2022-10-17 cs.CV cs.RO 79%

PedFormer: Pedestrian Behavior Prediction via Cross-Modal Attention Modulation and Gated Multitask Learning

Amir Rasouli, Iuliia Kotseruba

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

Comments 8 pages, 3 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.05840 2022-10-13 cs.CV 79%

LiveSeg: Unsupervised Multimodal Temporal Segmentation of Long Livestream Videos

Jielin Qiu, Franck Dernoncourt, Trung Bui, Zhaowen Wang, Ding Zhao, Hailin Jin

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.04331 2022-10-11 cs.CV 79%

Students taught by multimodal teachers are superior action recognizers

Gorjan Radevski, Dusan Grujicic, Matthew Blaschko, Marie-Francine Moens, Tinne Tuytelaars

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments Extended abstract accepted at the 2nd Ego4D Workshop @ ECCV 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.13371 2022-10-07 cs.CV 79%

FitCLIP: Refining Large-Scale Pretrained Image-Text Models for Zero-Shot Video Understanding Tasks

Santiago Castro, Fabian Caba Heilbron

专题命中 视频多模态 :image-text(title,abstract);分类 cs.CV

Comments Accepted at BMVC 2022. It includes the supplementary material. The margins and page size were modified to fit the arXiv ID stamp on the left side

详情

展开后加载摘要…

URL PDF HTML 收藏