arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 46477 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4778 篇

2403.19221 2024-03-29 cs.CV cs.AI 81%

Towards Multimodal Video Paragraph Captioning Models Robust to Missing Modality

Sishuo Chen, Lei Li, Shuhuai Ren, Rundong Gao, Yuanxin Liu, Xiaohan Bi, Xu Sun, Lu Hou

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Code available at https://github.com/lancopku/MR-VPC

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.15096 2024-02-26 cs.LG cs.CV cs.MM 81%

Multimodal Transformer With a Low-Computational-Cost Guarantee

Sungjin Park, Edward Choi

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM

Comments Accepted to ICASSP 2024 (5 pages)

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.08486 2024-02-22 cs.CL cs.CV 81%

Understanding Multimodal Procedural Knowledge by Sequencing Multimodal Instructional Manuals

Te-Lin Wu, Alex Spangher, Pegah Alipoormolabashi, Marjorie Freedman, Ralph Weischedel, Nanyun Peng

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments In Proceedings of the Conference of the 60th Annual Meeting of the Association for Computational Linguistics (ACL), 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.05653 2024-02-07 cs.CV cs.MM 81%

Semantic2Graph: Graph-based Multi-modal Feature Fusion for Action Segmentation in Videos

Junbin Zhang, Pei-Hsuan Tsai, Meng-Hsun Tsai

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.MM

Comments 13 pages, 3 figures, 9 tables. Published on Applied Intelligence

Journal ref Applied Intelligence(2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.16150 2024-02-05 cs.CV cs.AI cs.LG eess.IV 81%

Multimodal video and IMU kinematic dataset on daily life activities using affordable devices (VIDIMU)

Mario Martínez-Zarzuela, Javier González-Alonso, Míriam Antón-Rodríguez, Francisco J. Díaz-Pernas, Henning Müller, Cristina Simón-Martínez

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Journal ref Sci Data 10, 648 (2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.16076 2024-01-31 cs.CV cs.MM 81%

Find the Cliffhanger: Multi-Modal Trailerness in Soap Operas

Carlo Bretti, Pascal Mettes, Hendrik Vincent Koops, Daan Odijk, Nanne van Noord

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.MM

Comments MMM24

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.04965 2024-01-22 cs.CL cs.AI cs.LG 81%

MULTISCRIPT: Multimodal Script Learning for Supporting Open Domain Everyday Tasks

Jingyuan Qi, Minqian Liu, Ying Shen, Zhiyang Xu, Lifu Huang

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments Accepted by AAAI 2024. 11 pages, 9 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.09861 2024-01-19 cs.CV cs.AI 81%

Temporal Insight Enhancement: Mitigating Temporal Hallucination in Multimodal Large Language Models

Li Sun, Liuan Wang, Jun Sun, Takayuki Okatani

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 7 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.08581 2024-01-18 cs.CV cs.AI cs.LG 81%

Temporal Embeddings: Scalable Self-Supervised Temporal Representation Learning from Spatiotemporal Data for Multimodal Computer Vision

Yi Cao, Swetava Ganguli, Vipul Pandey

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Extended abstract accepted for presentation at BayLearn 2023. 3 pages, 7 figures. Abstract based on IEEE IGARSS 2023 research track paper: arXiv:2304.13143

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.06194 2024-01-15 cs.LG cs.AI cs.CL 81%

CrisisKAN: Knowledge-infused and Explainable Multimodal Attention Network for Crisis Event Classification

Shubham Gupta, Nandini Saini, Suman Kundu, Debasis Das

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.10763 2024-01-11 cs.CV cs.AI cs.LG eess.IV 81%

Actor-agnostic Multi-label Action Recognition with Multi-modal Query

Anindya Mondal, Sauradip Nag, Joaquin M Prada, Xiatian Zhu, Anjan Dutta

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments Published at the 2023 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW), Paris, France

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.03177 2024-01-09 cs.CV cs.CL 81%

Text-Video Retrieval via Variational Multi-Modal Hypergraph Networks

Qian Li, Lixin Su, Jiashu Zhao, Long Xia, Hengyi Cai, Suqi Cheng, Hengzhu Tang, Junfeng Wang, Dawei Yin

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.03040 2024-01-09 cs.LG cs.AI cs.CV cs.DC 81%

AccidentGPT: Large Multi-Modal Foundation Model for Traffic Accident Analysis

Kebin Wu, Wenbin Li, Xiaofei Xiao

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments 8 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.17117 2023-12-29 cs.CV cs.AI 81%

Grounding-Prompter: Prompting LLM with Multimodal Information for Temporal Sentence Grounding in Long Videos

Houlun Chen, Xin Wang, Hong Chen, Zihan Song, Jia Jia, Wenwu Zhu

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.01575 2023-12-05 cs.CL cs.CV 81%

A Challenging Multimodal Video Summary: Simultaneously Extracting and Generating Keyframe-Caption Pairs from Video

Keito Kudo, Haruki Nagasawa, Jun Suzuki, Nobuyuki Shimizu

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.10899 2023-11-22 cs.CV cs.CL cs.LG 81%

Extraction and Summarization of Explicit Video Content using Multi-Modal Deep Learning

Shaunak Joshi, Raghav Gaggar

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.CL

Comments 8 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.11845 2023-09-27 cs.SD cs.LG cs.MM eess.AS 81%

TMac: Temporal Multi-Modal Graph Learning for Acoustic Event Classification

Meng Liu, Ke Liang, Dayu Hu, Hao Yu, Yue Liu, Lingyuan Meng, Wenxuan Tu, Sihang Zhou, Xinwang Liu

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.MM、eess.AS

Comments This work has been accepted by ACM MM 2023 for publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.05880 2023-09-11 cs.CL cs.AI 81%

TikTalk: A Video-Based Dialogue Dataset for Multi-Modal Chitchat in Real World

Hongpeng Lin, Ludan Ruan, Wenke Xia, Peiyu Liu, Jingyuan Wen, Yixin Xu, Di Hu, Ruihua Song, Wayne Xin Zhao, Qin Jin, Zhiwu Lu

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CL、cs.AI

Comments Accepted to ACM Multimedia 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.14395 2023-08-29 cs.MM cs.CV 81%

UMMAFormer: A Universal Multimodal-adaptive Transformer Framework for Temporal Forgery Localization

Rui Zhang, Hongxia Wang, Mingshan Du, Hanqing Liu, Yang Zhou, Qiang Zeng

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM

Comments 11 pages, 8 figures, 66 references. This paper has been accepted for ACM MM 2023

Journal ref Proceedings of the 31st ACM International Conference on Multimedia (MM '23), October 29-November 3, 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.14160 2023-08-29 cs.CV cs.AI 81%

A Unified Transformer-based Network for multimodal Emotion Recognition

Kamran Ali, Charles E. Hughes

专题命中 视频多模态 :multimodal(title);multi-modal(abstract);分类 cs.CV、cs.AI

Comments 12 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.13309 2023-08-21 cs.CV cs.AI 81%

History Aware Multimodal Transformer for Vision-and-Language Navigation

Shizhe Chen, Pierre-Louis Guhur, Cordelia Schmid, Ivan Laptev

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted in NeurIPS 2021; project page at https://cshizhe.github.io/projects/vln_hamt.html; corrected a typo

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.00732 2023-08-14 cs.IR cs.AI cs.CL 81%

Kuaipedia: a Large-scale Multi-modal Short-video Encyclopedia

Haojie Pan, Zepeng Zhai, Yuzhou Zhang, Ruiji Fu, Ming Liu, Yangqiu Song, Zhongyuan Wang, Bing Qin

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.02469 2023-08-01 cs.CV cs.CL 81%

What Matters in Training a GPT4-Style Language Model with Multimodal Inputs?

Yan Zeng, Hanbo Zhang, Jiani Zheng, Jiangnan Xia, Guoqiang Wei, Yang Wei, Yuchen Zhang, Tao Kong

专题命中 视频多模态 :multimodal(title);multi-modal(abstract);分类 cs.CV、cs.CL

Comments 32 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.06385 2023-07-20 cs.CV cs.LG cs.SD eess.AS 81%

Temporal Label-Refinement for Weakly-Supervised Audio-Visual Event Localization

Kalyan Ramakrishnan

专题命中 视频多模态 :audio-visual(title,abstract);分类 cs.CV、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.00701 2023-06-26 astro-ph.IM astro-ph.CO cs.AI cs.CV cs.LG gr-qc 81%

DeepGraviLens: a Multi-Modal Architecture for Classifying Gravitational Lensing Data

Nicolò Oreste Pinciroli Vago, Piero Fraternali

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments This preprint has not undergone peer review or any post-submission improvements or corrections. The Version of Record of this article is published in Neural Computing and Applications, and is available online at https://doi.org/10.1007/s00521-023-08766-9

Journal ref Neural Comput & Applic (2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.12647 2023-06-08 cs.CV cs.AI 81%

Cross-Modal Causal Relational Reasoning for Event-Level Visual Question Answering

Yang Liu, Guanbin Li, Liang Lin

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV、cs.AI

Comments 17 pages, 9 figures. This work has been accepted by IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI). The datasets, code and models are available at https://github.com/HCPLab-SYSU/CMCIR

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.00355 2023-05-05 cs.CV cs.AI 81%

MH-DETR: Video Moment and Highlight Detection with Cross-modal Transformer

Yifang Xu, Yunzhuo Sun, Yang Li, Yilei Shi, Xiaoxiang Zhu, Sidan Du

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.01476 2023-05-03 cs.SD cs.MM eess.AS 81%

Deep Learning Based Multimodal with Two-phase Training Strategy for Daily Life Video Classification

Lam Pham, Trang Le, Cam Le, Dat Ngo, Weissenfeld Axel, Alexander Schindler

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.MM、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.07775 2023-04-28 cs.CV cs.MM 81%

Robust Cross-Modal Knowledge Distillation for Unconstrained Videos

Wenke Xia, Xingjian Li, Andong Deng, Haoyi Xiong, Dejing Dou, Di Hu

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.14369 2023-03-28 cs.CV cs.MM 81%

Video-Text as Game Players: Hierarchical Banzhaf Interaction for Cross-Modal Representation Learning

Peng Jin, Jinfa Huang, Pengfei Xiong, Shangxuan Tian, Chang Liu, Xiangyang Ji, Li Yuan, Jie Chen

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV、cs.MM

Comments CVPR 2023 Highlight

详情

展开后加载摘要…

URL PDF HTML 收藏