arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4729 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4729 篇

2410.03538 2024-10-22 cs.IR cs.AI cs.CV 84%

Dreaming User Multimodal Representation Guided by The Platonic Representation Hypothesis for Micro-Video Recommendation

Chengzhi Lin, Hezheng Lin, Shuchang Liu, Cangguang Ruan, LingJing Xu, Dezhao Yang, Chuyuan Wang, Yongqi Liu

专题命中 视频多模态 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.AI

Comments 4 Figure; 2 Table

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.14008 2024-08-27 cs.CV cs.AI 84%

LMM-VQA: Advancing Video Quality Assessment with Large Multimodal Models

Qihang Ge, Wei Sun, Yu Zhang, Yunhao Li, Zhongpeng Ji, Fengyu Sun, Shangling Jui, Xiongkuo Min, Guangtao Zhai

专题命中 视频多模态 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.12002 2024-07-18 cs.MM cs.CV 84%

A Multimodal Transformer for Live Streaming Highlight Prediction

Jiaxin Deng, Shiyao Wang, Dong Shen, Liqin Zhao, Fan Yang, Guorui Zhou, Gaofeng Meng

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.MM

Comments Accepted at ICME 2024 as poster presentation. arXiv admin note: text overlap with arXiv:2306.14392

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.07798 2024-06-11 cs.CV cs.AI 84%

FreeVA: Offline MLLM as Training-Free Video Assistant

Wenhao Wu

专题命中 视频多模态 :MLLM(title,abstract);multimodal(abstract);分类 cs.CV、cs.AI

Comments Preprint. Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.01174 2024-05-24 cs.CV cs.MM 84%

SpikeMba: Multi-Modal Spiking Saliency Mamba for Temporal Video Grounding

Wenrui Li, Xiaopeng Hong, Ruiqin Xiong, Xiaopeng Fan

专题命中 视频多模态 :multi-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.15494 2023-09-28 cs.CV cs.CL 84%

VideoAdviser: Video Knowledge Distillation for Multimodal Transfer Learning

Yanan Wang, Donghuo Zeng, Shinya Wada, Satoshi Kurihara

专题命中 视频多模态 :multimodal(title,abstract);audio-visual(abstract);分类 cs.CV、cs.CL

Comments Accepted by IEEE Access

Journal ref in IEEE Access, vol. 11, pp. 51229-51240, 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.05039 2023-07-04 cs.LG cs.AI cs.CV stat.ML 84%

Active Acquisition for Multimodal Temporal Data: A Challenging Decision-Making Task

Jannik Kossen, Cătălina Cangea, Eszter Vértes, Andrew Jaegle, Viorica Patraucean, Ira Ktena, Nenad Tomasev, Danielle Belgrave

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments Published in Transactions on Machine Learning Research. Previous version accepted to Foundation Models for Decision Making Workshop at NeurIPS 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.09146 2023-06-13 cs.CV cs.MM 84%

A Unified Multimodal De- and Re-coupling Framework for RGB-D Motion Recognition

Benjia Zhou, Pichao Wang, Jun Wan, Yanyan Liang, Fan Wang

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.MM

Comments Accepted to TPAMI 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.07214 2023-05-15 cs.CV cs.AI 84%

MMG-Ego4D: Multi-Modal Generalization in Egocentric Action Recognition

Xinyu Gong, Sreyas Mohan, Naina Dhingra, Jean-Charles Bazin, Yilei Li, Zhangyang Wang, Rakesh Ranjan

专题命中 视频多模态 :multi-modal(title);multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted to CVPR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.02080 2023-04-06 cs.CV cs.CL 84%

Scalable and Accurate Self-supervised Multimodal Representation Learning without Aligned Video and Text Data

Vladislav Lialin, Stephen Rawls, David Chan, Shalini Ghosh, Anna Rumshisky, Wael Hamza

专题命中 视频多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.CL

Journal ref 2023 IEEE/CVF Winter Conference on Applications of Computer Vision Workshops (WACVW)

详情

展开后加载摘要…

URL PDF HTML 收藏
2009.10434 2023-04-06 cs.CV cs.MM 84%

Frame-wise Cross-modal Matching for Video Moment Retrieval

Haoyu Tang, Jihua Zhu, Meng Liu, Zan Gao, Zhiyong Cheng

专题命中 视频多模态 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.MM

Comments 12 pages; accepted by IEEE TMM

Journal ref IEEE Transactions on Multimedia 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.11357 2023-01-30 cs.CV cs.CL 84%

Multimodal Event Transformer for Image-guided Story Ending Generation

Yucheng Zhou, Guodong Long

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.CL

Comments EACL 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.11544 2022-11-03 cs.CV cs.AI 84%

Rethinking Multi-Modal Alignment in Video Question Answering from Feature and Sample Perspectives

Shaoning Xiao, Long Chen, Kaifeng Gao, Zhao Wang, Yi Yang, Zhimeng Zhang, Jun Xiao

专题命中 视频多模态 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.08090 2022-08-18 cs.CV cs.MM 84%

Progressive Cross-modal Knowledge Distillation for Human Action Recognition

Jianyuan Ni, Anne H. H. Ngu, Yan Yan

专题命中 视频多模态 :cross-modal(title);multimodal(abstract);multi-modal(abstract);分类 cs.CV、cs.MM

Comments ACM MM 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.06085 2022-04-12 cs.CV cs.CL 84%

On Pursuit of Designing Multi-modal Transformer for Video Grounding

Meng Cao, Long Chen, Mike Zheng Shou, Can Zhang, Yuexian Zou

专题命中 视频多模态 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.CL

Comments Accepted by Conference on Empirical Methods in Natural Language Processing (EMNLP 2021, Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.02566 2022-04-07 cs.CL cs.MM 84%

Modeling Temporal-Modal Entity Graph for Procedural Multimodal Machine Comprehension

Huibin Zhang, Zhengkun Zhang, Yao Zhang, Jun Wang, Yufan Li, Ning jiang, Xin wei, Zhenglu Yang

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL、cs.MM

Comments Accepted by ACL-2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.07175 2021-05-18 cs.CV cs.MM 84%

Cross-Modal Progressive Comprehension for Referring Segmentation

Si Liu, Tianrui Hui, Shaofei Huang, Yunchao Wei, Bo Li, Guanbin Li

专题命中 视频多模态 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.MM

Comments Accepted by TPAMI 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2007.03848 2021-03-04 cs.CV cs.CL 84%

Dynamic Graph Representation Learning for Video Dialog via Multi-Modal Shuffled Transformers

Shijie Geng, Peng Gao, Moitreya Chatterjee, Chiori Hori, Jonathan Le Roux, Yongfeng Zhang, Hongsheng Li, Anoop Cherian

专题命中 视频多模态 :multi-modal(title,abstract);audio-visual(abstract);分类 cs.CV、cs.CL

Comments Accepted at AAAI 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
1907.03049 2020-02-18 cs.CV cs.CL 84%

Video Question Generation via Cross-Modal Self-Attention Networks Learning

Yu-Siang Wang, Hung-Ting Su, Chen-Hsi Chang, Zhe-Yu Liu, Winston H. Hsu

专题命中 视频多模态 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.CL

Comments Accepted by ICASSP 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
1901.04268 2019-05-23 cs.IR cs.CV cs.MM 84%

Learning Shared Semantic Space with Correlation Alignment for Cross-modal Event Retrieval

Zhenguo Yang, Zehang Lin, Peipei Kang, Jianming Lv, Qing Li, Wenyin Liu

专题命中 视频多模态 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.MM

Comments 22 pages, submitted to ACM Transactions on Multimedia Computing Communications and Applications(ACM TOMM)

详情

展开后加载摘要…

URL PDF HTML 收藏
1812.07023 2018-12-19 cs.CL cs.CV 84%

From FiLM to Video: Multi-turn Question Answering with Multi-modal Context

Dat Tien Nguyen, Shikhar Sharma, Hannes Schulz, Layla El Asri

专题命中 视频多模态 :multi-modal(title,abstract);audio-visual(abstract);分类 cs.CV、cs.CL

Comments Accepted for an Oral presentation at the DSTC7 workshop at AAAI 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1512.00818 2015-12-17 cs.CV cs.CL cs.LG 84%

Zero-Shot Event Detection by Multimodal Distributional Semantic Embedding of Videos

Mohamed Elhoseiny, Jingen Liu, Hui Cheng, Harpreet Sawhney, Ahmed Elgammal

专题命中 视频多模态 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.CL

Comments To appear in AAAI 2016

详情

展开后加载摘要…

URL PDF HTML 收藏
1308.1150 2013-08-07 cs.MM cs.CV 84%

Multimodal Approach for Video Surveillance Indexing and Retrieval

Ali Wali, Adel M. Alimi

专题命中 视频多模态 :multimodal(title,abstract);audio-visual(abstract);分类 cs.CV、cs.MM

Comments 7 pages

Journal ref Journal of Intelligent Computing, Volume: 1, Issue: 4 (December 2010), Page: 165-175

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.10783 2023-09-20 cs.CV cs.AI cs.CL 83%

Language as the Medium: Multimodal Video Classification through text only

Laura Hanu, Anita L. Verő, James Thewlis

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI;multimodal foundation model(comments)

Comments Accepted at "What is Next in Multimodal Foundation Models?" (MMFM) workshop at ICCV 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14854 2026-08-18 cs.CV 新提交 83%

Zero-MELO: Test-Time Evidence Calibration with Multimodal LLMs for Zero-Shot Micro-Gesture Recognition

Zero-MELO:基于多模态大语言模型的测试时证据校准用于零样本微手势识别

Chengyan Wang, Hanliang Xie, Yueyi Yang, Haoyu Chen

机构 * University of Oulu(奥卢大学) Peking University(北京大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 该研究针对多模态大语言模型在微手势识别中局部证据不足、分数偏差的瓶颈,提出Zero-MELO框架,结合树搜索、测试时校准与多线索融合,在iMiGUE和MA-52数据集上显著优于Qwen2.5-VL基线。

Comments Accepted by ACM MM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01045 2026-08-11 cs.AI 版本更新 83%

Med-CRAFT: An Information System for Explainable and Configurable Construction of Multimodal Medical QA Datasets

Med-CRAFT:通过知识图谱遍历自动构建可解释的多跳视频工作负载

Shenxi Liu, Kan Li, Mingyang Zhao, Yuhang Tian, Bin Li

机构 * Beijing Institute of Technology(北京理工大学) The Hong Kong Polytechnic University(香港理工大学) Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究所)

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract_cn);分类 cs.AI

AI总结 Med-CRAFT通过知识图谱遍历自动构建可解释的多跳视频工作负载,生成具有细粒度时间选择性和多跳逻辑复杂性的医疗视频推理基准。

Comments 23 pages, 4 figures, 10 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.02324 2026-08-04 cs.CV 新提交 83%

Implicit Neural Representations for Multimodal Longitudinal Image Imputation and Interpolation

用于多模态纵向图像插补与插值的隐式神经表示

Sina Wendrich, Lukas Förner, Zoe Reinke, Kartikay Tehlan, Ansgar Berlis, Michael Frühwald, Matthias Wagner, Thomas Wendler

机构 * University Hospital Augsburg(奥格斯堡大学医院) University of Augsburg(奥格斯堡大学) Technical University of Munich(慕尼黑工业大学) Bavarian Cancer Research Center (BZKF)(巴伐利亚癌症研究中心) Swabian Children’s Cancer Center(施瓦本儿童癌症中心)

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 该研究针对临床纵向MRI数据缺失等问题,提出条件隐式神经表示模型,经儿科脑肿瘤数据验证,可显著改进插值效果,置信度估计可靠。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03168 2026-07-28 cs.CV cs.LG 版本更新 83%

Farm-LightSeek: An Edge-centric Multimodal Agricultural IoT Data Analytics Framework with Lightweight LLMs

Farm-LightSeek:一种以边缘为中心的多模态农业物联网数据分析框架,集成轻量级语言模型

Dawen Jiang, Zhishu Shen, Qiushi Zheng, Tiehua Zhang, Wei Xiang, Jiong Jin

机构 * School of Computer Science and Artificial Intelligence, Wuhan University of Technology(武汉理工大学计算机科学与人工智能学院) School of Engineering, Swinburne University of Technology(斯威本科技大学工程学院) School of Computer Science and Technology, Tongji University(同济大学计算机科学与技术学院) School of Computing, Engineering and Mathematical Sciences, La Trobe University(拉筹伯大学计算、工程与数学科学学院)

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 针对智能农业面临的挑战,提出Farm-LightSeek框架,将大语言模型与边缘计算集成。通过传感器收集多源数据,在边缘节点进行跨模态推理等。创新包括闭环架构等,实验表明该框架在关键任务中性能可靠,推动了智能实时农业及两者深度集成。

Comments Accepted by IEEE Internet of Things Magazine

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.20981 2026-07-24 cs.AI 新提交 83%

Beyond Independent Optimization: Compression, MoE Routing, and Quantization Interactions in Multimodal Edge Intelligence

超越独立优化:多模态边缘智能中的压缩、混合专家路由和量化交互

Jay Gor, Karm Dave, Akshita Abrol, Rajesh Gupta, Sudeep Tanwar, Zhengkui Wang

机构 * Nirma University(尼玛大学) Singapore Institute of Technology(新加坡理工学院) Marwadi University(马尔瓦迪大学)

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

AI总结 研究多模态边缘智能中高效推理受多种因素限制,回顾相关模型进展,指出技术间相互影响不能独立优化,介绍关键设计权衡,引入视频MoE模型诊断方法,强调多方面开放研究方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.17712 2026-07-21 cs.AI 新提交 83%

Learning to Detect Cross-Modal Negation: An Analysis of Latent Representations and an Attention-Based Solution

学习检测跨模态否定:潜在表示分析与基于注意力的解决方案

Ali AbuSaleh, Leon Hammerla, Alexander Mehler

专题命中 视频多模态 :cross-modal(title,abstract);multimodal(abstract);分类 cs.AI

AI总结 研究跨模态否定检测难题,提出新型跨模态注意力架构,分析发现文本与视觉否定的不对称性,结合自监督视频表示推进时间否定建模,为多模态系统学习语义对齐表示提供新方法。

Comments This manuscript is an accepted version of the article (published at ICNLP2026). Published in IEEE Xplore, DOI:10.1109/ICNLP69856.2026.11527861 document: https://ieeexplore.ieee.org/abstract/document/11527861

Journal ref 2026 8th International Conference on Natural Language Processing (ICNLP), Xi'an, China, 2026, pp. 613-622

详情

展开后加载摘要…

URL PDF HTML 收藏