arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4735 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4735 篇

2501.01960 2025-01-07 cs.CV cs.AI cs.GR cs.LG 81%

GAF-FusionNet: Multimodal ECG Analysis via Gramian Angular Fields and Split Attention

Jiahao Qin, Feng Liu

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 14 pages, 1 figure, accepted by ICONIP 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.11621 2024-12-17 cs.CV cs.MM 81%

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting

Muhammet Furkan Ilaslan, Ali Koksal, Kevin Qinhong Lin, Burak Satar, Mike Zheng Shou, Qianli Xu

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM

Comments Accepted for The 39th Annual AAAI Conference on Artificial Intelligence 2025 in Main Track, 19 pages, 24 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.10360 2024-12-16 cs.CV cs.AI 81%

Apollo: An Exploration of Video Understanding in Large Multimodal Models

Orr Zohar, Xiaohan Wang, Yann Dubois, Nikhil Mehta, Tong Xiao, Philippe Hansen-Estruch, Licheng Yu, Xiaofang Wang, Felix Juefei-Xu, Ning Zhang, Serena Yeung-Levy, Xide Xia

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments https://apollo-lmms.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.18938 2024-12-04 cs.CV cs.AI 81%

From Seconds to Hours: Reviewing MultiModal Large Language Models on Comprehensive Long Video Understanding

Heqing Zou, Tianze Luo, Guiyang Xie, Victor, Zhang, Fengmao Lv, Guangcong Wang, Junyang Chen, Zhuochen Wang, Hansheng Zhang, Huaijian Zhang

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 11 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.13053 2024-11-21 cs.CV cs.AI cs.LG 81%

MEGL: Multimodal Explanation-Guided Learning

Yifei Zhang, Tianxu Jiang, Bo Pan, Jingyu Wang, Guangji Bai, Liang Zhao

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.16620 2024-11-13 cs.CV cs.CL 81%

OmAgent: A Multi-modal Agent Framework for Complex Video Understanding with Task Divide-and-Conquer

Lu Zhang, Tiancheng Zhao, Heting Ying, Yibo Ma, Kyusong Lee

专题命中 视频多模态 :multi-modal(title);multimodal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.05767 2024-10-28 cs.AI cs.CL cs.IR 81%

A Survey of Knowledge Graph Reasoning on Graph Types: Static, Dynamic, and Multimodal

Ke Liang, Lingyuan Meng, Meng Liu, Yue Liu, Wenxuan Tu, Siwei Wang, Sihang Zhou, Xinwang Liu, Fuchun Sun

专题命中 视频多模态 :multimodal(title);multi-modal(abstract);分类 cs.CL、cs.AI

Comments This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.16116 2024-10-22 astro-ph.SR astro-ph.IM cs.AI cs.CV 81%

Multimodal Flare Forecasting with Deep Learning

Grégoire Francisco, Sabrina Guastavino, Teresa Barata, João Fernandes, Dario Del Moro

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.14150 2024-10-21 cs.AI cs.CL 81%

Utilizing Large Language Models for Event Deconstruction to Enhance Multimodal Aspect-Based Sentiment Analysis

Xiaoyong Huang, Heli Sun, Qunshu Gao, Wenjie Huang, Ruichen Cao

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.09379 2024-10-15 cs.CV cs.AI 81%

Multi-granularity Contrastive Cross-modal Collaborative Generation for End-to-End Long-term Video Question Answering

Ting Yu, Kunhao Fu, Jian Zhang, Qingming Huang, Jun Yu

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV、cs.AI

Comments Transactions on Image Processing

Journal ref Transactions on Image Processing, vol. 33, pp. 3115-3129, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.07405 2024-10-11 cs.CV cs.AI 81%

Exploring Efficient Foundational Multi-modal Models for Video Summarization

Karan Samel, Apoorva Beedu, Nitish Sontakke, Irfan Essa

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments 11 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.05267 2024-10-08 cs.CL cs.CV 81%

Grounding Partially-Defined Events in Multimodal Data

Kate Sanders, Reno Kriz, David Etter, Hannah Recknor, Alexander Martin, Cameron Carpenter, Jingyang Lin, Benjamin Van Durme

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Preprint; 9 pages; 2024 EMNLP Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.16050 2024-10-04 cs.CV cs.CL 81%

Efficient Temporal Extrapolation of Multimodal Large Language Models with Temporal Grounding Bridge

Yuxuan Wang, Yueqian Wang, Pengfei Wu, Jianxin Liang, Dongyan Zhao, Yang Liu, Zilong Zheng

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments To appear at EMNLP 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.16099 2024-09-25 cs.CV cs.AI 81%

Neuromorphic Drone Detection: an Event-RGB Multimodal Approach

Gabriele Magrini, Federico Becattini, Pietro Pala, Alberto Del Bimbo, Antonio Porta

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted at NeVi Workshop at ECCV24

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.09446 2024-09-17 cs.CV cs.AI 81%

MulCPred: Learning Multi-modal Concepts for Explainable Pedestrian Action Prediction

Yan Feng, Alexander Carballo, Keisuke Fujii, Robin Karlsson, Ming Ding, Kazuya Takeda

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.21757 2024-09-13 cs.CV cs.MM 81%

Learning Video Context as Interleaved Multimodal Sequences

Kevin Qinghong Lin, Pengchuan Zhang, Difei Gao, Xide Xia, Joya Chen, Ziteng Gao, Jinheng Xie, Xuhong Xiao, Mike Zheng Shou

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM

Comments Accepted by ECCV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.01560 2024-09-04 cs.CV cs.AI 81%

Blocks as Probes: Dissecting Categorization Ability of Large Multimodal Models

Bin Fu, Qiyang Wan, Jialin Li, Ruiping Wang, Xilin Chen

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 39 pages, 28 figures, 4 tables. Accepted at The 35th British Machine Vision Conference (BMVC 2024). Project page at https://fubin29.github.io/Blocks-as-Probes/

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.10579 2024-09-04 cs.CL cs.AI 81%

DER-GCN: Dialogue and Event Relation-Aware Graph Convolutional Neural Network for Multimodal Dialogue Emotion Recognition

Wei Ai, Yuntao Shou, Tao Meng, Nan Yin, Keqin Li

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments 14 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.17428 2024-08-27 cs.CV cs.AI 81%

SigFormer: Sparse Signal-Guided Transformer for Multi-Modal Human Action Segmentation

Qi Liu, Xinchen Liu, Kun Liu, Xiaoyan Gu, Wu Liu

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.07981 2024-08-16 cs.CV cs.AI 81%

LLaVA-Surg: Towards Multimodal Surgical Assistant via Structured Surgical Video Learning

Jiajie Li, Garrett Skinner, Gene Yang, Brian R Quaranto, Steven D Schwaitzberg, Peter C W Kim, Jinjun Xiong

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.04388 2024-08-09 cs.MM cs.AI cs.IR 81%

MM-Forecast: A Multimodal Approach to Temporal Event Forecasting with Large Language Models

Haoxuan Li, Zhengmao Yang, Yunshan Ma, Yi Bin, Yang Yang, Tat-Seng Chua

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.07895 2024-07-30 cs.CV cs.CL cs.LG 81%

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Feng Li, Renrui Zhang, Hao Zhang, Yuanhan Zhang, Bo Li, Wei Li, Zejun Ma, Chunyuan Li

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Project Page: https://llava-vl.github.io/blog/2024-06-16-llava-next-interleave/

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.06157 2024-07-09 cs.CV cs.AI 81%

Temporal Grounding of Activities using Multimodal Large Language Models

Young Chol Song

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.04923 2024-07-09 cs.CV cs.CL 81%

OmChat: A Recipe to Train Multimodal Language Models with Strong Long Context and Video Understanding

Tiancheng Zhao, Qianqian Zhang, Kyusong Lee, Peng Liu, Lu Zhang, Chunxin Fang, Jiajia Liao, Kelei Jiang, Yibo Ma, Ruochen Xu

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments 14 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.04697 2024-07-08 cs.CV cs.MM 81%

VCoME: Verbal Video Composition with Multimodal Editing Effects

Weibo Gong, Xiaojie Jin, Xin Li, Dongliang He, Xinglong Wu

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.18139 2024-06-27 cs.CL cs.CV 81%

LOOK-M: Look-Once Optimization in KV Cache for Efficient Multimodal Long-Context Inference

Zhongwei Wan, Ziang Wu, Che Liu, Jinfa Huang, Zhihong Zhu, Peng Jin, Longyue Wang, Li Yuan

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.04873 2024-05-22 cs.CV cs.AI 81%

Multimodal Prototype-Enhanced Network for Few-Shot Action Recognition

Xinzhe Ni, Yong Liu, Hao Wen, Yatai Ji, Jing Xiao, Yujiu Yang

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted by ICMR 2024 (oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.05130 2024-05-09 cs.CV cs.MM 81%

Multi-scale Bottleneck Transformer for Weakly Supervised Multimodal Violence Detection

Shengyang Sun, Xiaojin Gong

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM

Comments Accepted by ICME 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.16557 2024-04-26 cs.CV cs.AI 81%

Energy-Latency Manipulation of Multi-modal Large Language Models via Verbose Samples

Kuofeng Gao, Jindong Gu, Yang Bai, Shu-Tao Xia, Philip Torr, Wei Liu, Zhifeng Li

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments arXiv admin note: substantial text overlap with arXiv:2401.11170

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.07484 2024-04-12 cs.MM cs.AI 81%

Multimodal Emotion Recognition by Fusing Video Semantic in MOOC Learning Scenarios

Yuan Zhang, Xiaomei Tao, Hanxu Ai, Tao Chen, Yanling Gan

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏