arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4749 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4749 篇

2404.16557 2024-04-26 cs.CV cs.AI 81%

Energy-Latency Manipulation of Multi-modal Large Language Models via Verbose Samples

Kuofeng Gao, Jindong Gu, Yang Bai, Shu-Tao Xia, Philip Torr, Wei Liu, Zhifeng Li

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments arXiv admin note: substantial text overlap with arXiv:2401.11170

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.07484 2024-04-12 cs.MM cs.AI 81%

Multimodal Emotion Recognition by Fusing Video Semantic in MOOC Learning Scenarios

Yuan Zhang, Xiaomei Tao, Hanxu Ai, Tao Chen, Yanling Gan

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.04763 2024-04-09 cs.CV cs.AI 81%

GenEARL: A Training-Free Generative Framework for Multimodal Event Argument Role Labeling

Hritik Bansal, Po-Nien Kung, P. Jeffrey Brantingham, Kai-Wei Chang, Nanyun Peng

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 20 pages, 15 Figures, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.01258 2024-04-03 cs.CV cs.AI 81%

Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward

Ruohong Zhang, Liangke Gui, Zhiqing Sun, Yihao Feng, Keyang Xu, Yuanhan Zhang, Di Fu, Chunyuan Li, Alexander Hauptmann, Yonatan Bisk, Yiming Yang

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.19221 2024-03-29 cs.CV cs.AI 81%

Towards Multimodal Video Paragraph Captioning Models Robust to Missing Modality

Sishuo Chen, Lei Li, Shuhuai Ren, Rundong Gao, Yuanxin Liu, Xiaohan Bi, Xu Sun, Lu Hou

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Code available at https://github.com/lancopku/MR-VPC

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.15096 2024-02-26 cs.LG cs.CV cs.MM 81%

Multimodal Transformer With a Low-Computational-Cost Guarantee

Sungjin Park, Edward Choi

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM

Comments Accepted to ICASSP 2024 (5 pages)

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.08486 2024-02-22 cs.CL cs.CV 81%

Understanding Multimodal Procedural Knowledge by Sequencing Multimodal Instructional Manuals

Te-Lin Wu, Alex Spangher, Pegah Alipoormolabashi, Marjorie Freedman, Ralph Weischedel, Nanyun Peng

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments In Proceedings of the Conference of the 60th Annual Meeting of the Association for Computational Linguistics (ACL), 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.05653 2024-02-07 cs.CV cs.MM 81%

Semantic2Graph: Graph-based Multi-modal Feature Fusion for Action Segmentation in Videos

Junbin Zhang, Pei-Hsuan Tsai, Meng-Hsun Tsai

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.MM

Comments 13 pages, 3 figures, 9 tables. Published on Applied Intelligence

Journal ref Applied Intelligence(2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.16150 2024-02-05 cs.CV cs.AI cs.LG eess.IV 81%

Multimodal video and IMU kinematic dataset on daily life activities using affordable devices (VIDIMU)

Mario Martínez-Zarzuela, Javier González-Alonso, Míriam Antón-Rodríguez, Francisco J. Díaz-Pernas, Henning Müller, Cristina Simón-Martínez

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Journal ref Sci Data 10, 648 (2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.16076 2024-01-31 cs.CV cs.MM 81%

Find the Cliffhanger: Multi-Modal Trailerness in Soap Operas

Carlo Bretti, Pascal Mettes, Hendrik Vincent Koops, Daan Odijk, Nanne van Noord

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.MM

Comments MMM24

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.04965 2024-01-22 cs.CL cs.AI cs.LG 81%

MULTISCRIPT: Multimodal Script Learning for Supporting Open Domain Everyday Tasks

Jingyuan Qi, Minqian Liu, Ying Shen, Zhiyang Xu, Lifu Huang

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments Accepted by AAAI 2024. 11 pages, 9 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.09861 2024-01-19 cs.CV cs.AI 81%

Temporal Insight Enhancement: Mitigating Temporal Hallucination in Multimodal Large Language Models

Li Sun, Liuan Wang, Jun Sun, Takayuki Okatani

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 7 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.08581 2024-01-18 cs.CV cs.AI cs.LG 81%

Temporal Embeddings: Scalable Self-Supervised Temporal Representation Learning from Spatiotemporal Data for Multimodal Computer Vision

Yi Cao, Swetava Ganguli, Vipul Pandey

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Extended abstract accepted for presentation at BayLearn 2023. 3 pages, 7 figures. Abstract based on IEEE IGARSS 2023 research track paper: arXiv:2304.13143

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.06194 2024-01-15 cs.LG cs.AI cs.CL 81%

CrisisKAN: Knowledge-infused and Explainable Multimodal Attention Network for Crisis Event Classification

Shubham Gupta, Nandini Saini, Suman Kundu, Debasis Das

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.10763 2024-01-11 cs.CV cs.AI cs.LG eess.IV 81%

Actor-agnostic Multi-label Action Recognition with Multi-modal Query

Anindya Mondal, Sauradip Nag, Joaquin M Prada, Xiatian Zhu, Anjan Dutta

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments Published at the 2023 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW), Paris, France

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.03177 2024-01-09 cs.CV cs.CL 81%

Text-Video Retrieval via Variational Multi-Modal Hypergraph Networks

Qian Li, Lixin Su, Jiashu Zhao, Long Xia, Hengyi Cai, Suqi Cheng, Hengzhu Tang, Junfeng Wang, Dawei Yin

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.03040 2024-01-09 cs.LG cs.AI cs.CV cs.DC 81%

AccidentGPT: Large Multi-Modal Foundation Model for Traffic Accident Analysis

Kebin Wu, Wenbin Li, Xiaofei Xiao

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments 8 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.17117 2023-12-29 cs.CV cs.AI 81%

Grounding-Prompter: Prompting LLM with Multimodal Information for Temporal Sentence Grounding in Long Videos

Houlun Chen, Xin Wang, Hong Chen, Zihan Song, Jia Jia, Wenwu Zhu

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.01575 2023-12-05 cs.CL cs.CV 81%

A Challenging Multimodal Video Summary: Simultaneously Extracting and Generating Keyframe-Caption Pairs from Video

Keito Kudo, Haruki Nagasawa, Jun Suzuki, Nobuyuki Shimizu

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.10899 2023-11-22 cs.CV cs.CL cs.LG 81%

Extraction and Summarization of Explicit Video Content using Multi-Modal Deep Learning

Shaunak Joshi, Raghav Gaggar

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.CL

Comments 8 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.11845 2023-09-27 cs.SD cs.LG cs.MM eess.AS 81%

TMac: Temporal Multi-Modal Graph Learning for Acoustic Event Classification

Meng Liu, Ke Liang, Dayu Hu, Hao Yu, Yue Liu, Lingyuan Meng, Wenxuan Tu, Sihang Zhou, Xinwang Liu

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.MM、eess.AS

Comments This work has been accepted by ACM MM 2023 for publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.05880 2023-09-11 cs.CL cs.AI 81%

TikTalk: A Video-Based Dialogue Dataset for Multi-Modal Chitchat in Real World

Hongpeng Lin, Ludan Ruan, Wenke Xia, Peiyu Liu, Jingyuan Wen, Yixin Xu, Di Hu, Ruihua Song, Wayne Xin Zhao, Qin Jin, Zhiwu Lu

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CL、cs.AI

Comments Accepted to ACM Multimedia 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.14395 2023-08-29 cs.MM cs.CV 81%

UMMAFormer: A Universal Multimodal-adaptive Transformer Framework for Temporal Forgery Localization

Rui Zhang, Hongxia Wang, Mingshan Du, Hanqing Liu, Yang Zhou, Qiang Zeng

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM

Comments 11 pages, 8 figures, 66 references. This paper has been accepted for ACM MM 2023

Journal ref Proceedings of the 31st ACM International Conference on Multimedia (MM '23), October 29-November 3, 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.14160 2023-08-29 cs.CV cs.AI 81%

A Unified Transformer-based Network for multimodal Emotion Recognition

Kamran Ali, Charles E. Hughes

专题命中 视频多模态 :multimodal(title);multi-modal(abstract);分类 cs.CV、cs.AI

Comments 12 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.13309 2023-08-21 cs.CV cs.AI 81%

History Aware Multimodal Transformer for Vision-and-Language Navigation

Shizhe Chen, Pierre-Louis Guhur, Cordelia Schmid, Ivan Laptev

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted in NeurIPS 2021; project page at https://cshizhe.github.io/projects/vln_hamt.html; corrected a typo

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.00732 2023-08-14 cs.IR cs.AI cs.CL 81%

Kuaipedia: a Large-scale Multi-modal Short-video Encyclopedia

Haojie Pan, Zepeng Zhai, Yuzhou Zhang, Ruiji Fu, Ming Liu, Yangqiu Song, Zhongyuan Wang, Bing Qin

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.02469 2023-08-01 cs.CV cs.CL 81%

What Matters in Training a GPT4-Style Language Model with Multimodal Inputs?

Yan Zeng, Hanbo Zhang, Jiani Zheng, Jiangnan Xia, Guoqiang Wei, Yang Wei, Yuchen Zhang, Tao Kong

专题命中 视频多模态 :multimodal(title);multi-modal(abstract);分类 cs.CV、cs.CL

Comments 32 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.06385 2023-07-20 cs.CV cs.LG cs.SD eess.AS 81%

Temporal Label-Refinement for Weakly-Supervised Audio-Visual Event Localization

Kalyan Ramakrishnan

专题命中 视频多模态 :audio-visual(title,abstract);分类 cs.CV、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.00701 2023-06-26 astro-ph.IM astro-ph.CO cs.AI cs.CV cs.LG gr-qc 81%

DeepGraviLens: a Multi-Modal Architecture for Classifying Gravitational Lensing Data

Nicolò Oreste Pinciroli Vago, Piero Fraternali

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments This preprint has not undergone peer review or any post-submission improvements or corrections. The Version of Record of this article is published in Neural Computing and Applications, and is available online at https://doi.org/10.1007/s00521-023-08766-9

Journal ref Neural Comput & Applic (2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.12647 2023-06-08 cs.CV cs.AI 81%

Cross-Modal Causal Relational Reasoning for Event-Level Visual Question Answering

Yang Liu, Guanbin Li, Liang Lin

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV、cs.AI

Comments 17 pages, 9 figures. This work has been accepted by IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI). The datasets, code and models are available at https://github.com/HCPLab-SYSU/CMCIR

详情

展开后加载摘要…

URL PDF HTML 收藏