arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4749 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4749 篇

2404.07610 2024-04-12 cs.CV 79%

Do You Remember? Dense Video Captioning with Cross-Modal Memory Retrieval

Minkuk Kim, Hyeon Bae Kim, Jinyoung Moon, Jinwoo Choi, Seong Tae Kim

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

Comments CVPR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.07540 2024-04-09 cs.LG cs.CV q-bio.QM 79%

Tensor-based Multimodal Learning for Prediction of Pulmonary Arterial Wedge Pressure from Cardiac MRI

Prasun C. Tripathi, Mohammod N. I. Suvon, Lawrence Schobs, Shuo Zhou, Samer Alabed, Andrew J. Swift, Haiping Lu

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.05698 2024-04-05 cs.CV 79%

Mirasol3B: A Multimodal Autoregressive model for time-aligned and contextual modalities

AJ Piergiovanni, Isaac Noble, Dahun Kim, Michael S. Ryoo, Victor Gomes, Anelia Angelova

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments CVPR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.19811 2024-04-01 cs.CV 79%

X-MIC: Cross-Modal Instance Conditioning for Egocentric Action Generalization

Anna Kukleva, Fadime Sener, Edoardo Remelli, Bugra Tekin, Eric Sauser, Bernt Schiele, Shugao Ma

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

Comments CVPR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.12328 2024-03-26 cs.LG cs.AI 79%

A survey on knowledge-enhanced multimodal learning

Maria Lymperaiou, Giorgos Stamou

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.14598 2024-03-22 cs.CV 79%

PSALM: Pixelwise SegmentAtion with Large Multi-Modal Model

Zheng Zhang, Yeyao Ma, Enming Zhang, Xiang Bai

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.14082 2024-03-22 cs.CV 79%

EventDance: Unsupervised Source-free Cross-modal Adaptation for Event-based Object Recognition

Xu Zheng, Lin Wang

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

Comments Accepted to CVPR2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.13507 2024-03-22 cs.CV 79%

FMM-Attack: A Flow-based Multi-modal Adversarial Attack on Video-based LLMs

Jinmin Li, Kuofeng Gao, Yang Bai, Jingyun Zhang, Shu-tao Xia, Yisen Wang

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.16198 2024-03-08 cs.CV cs.LG 79%

Multi-modal learning for geospatial vegetation forecasting

Vitus Benson, Claire Robin, Christian Requena-Mesa, Lazaro Alonso, Nuno Carvalhais, José Cortés, Zhihan Gao, Nora Linscheid, Mélanie Weynants, Markus Reichstein

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted at CVPR 2024. We provide open source code and pre-trained weights to reproduce our experimental results under https://github.com/vitusbenson/greenearthnet

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.02483 2024-03-07 cs.CV 79%

EtC: Temporal Boundary Expand then Clarify for Weakly Supervised Video Grounding with Multimodal Large Language Model

Guozhang Li, Xinpeng Ding, De Cheng, Jie Li, Nannan Wang, Xinbo Gao

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.11403 2024-03-05 cs.AI 79%

An Empirical Evaluation of Neural and Neuro-symbolic Approaches to Real-time Multimodal Complex Event Detection

Liying Han, Mani B. Srivastava

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.00366 2024-03-04 cs.CY cs.CV 79%

Exploring the dynamic interplay of cognitive load and emotional arousal by using multimodal measurements: Correlation of pupil diameter and emotional arousal in emotionally engaging tasks

C. Kosel, S. Michel, T. Seidel, M. Foerster

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments The first two authors contributed equally to the manuscript

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.13270 2024-02-22 physics.ao-ph cs.AI cs.LG physics.data-an 79%

Global Tropical Cyclone Intensity Forecasting with Multi-modal Multi-scale Causal Autoregressive Model

Xinyu Wang, Kang Chen, Lei Liu, Tao Han, Bin Li, Lei Bai

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.13146 2024-02-21 cs.CV 79%

OLViT: Multi-Modal State Tracking via Attention-Based Embeddings for Video-Grounded Dialog

Adnen Abdessaied, Manuel von Hochmeister, Andreas Bulling

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments COLING 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.12845 2024-02-21 cs.AI cs.GT 79%

MORE-3S:Multimodal-based Offline Reinforcement Learning with Shared Semantic Spaces

Tianyu Zheng, Ge Zhang, Xingwei Qu, Ming Kuang, Stephen W. Huang, Zhaofeng He

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.11875 2024-02-20 cs.CL 79%

M2K-VDG: Model-Adaptive Multimodal Knowledge Anchor Enhanced Video-grounded Dialogue Generation

Hongcheng Liu, Pingjie Wang, Yu Wang, Yanfeng Wang

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.08607 2024-02-16 cs.CY cs.CV cs.LG 79%

Monitoring of Urban Changes with multi-modal Sentinel 1 and 2 Data in Mariupol, Ukraine, in 2022/23

Georg Zitzlsberger, Michal Podhoranyi

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted for publication in IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.00965 2024-02-05 cs.LG cs.CV eess.SP 79%

Multi-Modal Machine Learning Framework for Automated Seizure Detection in Laboratory Rats

Aaron Mullen, Samuel E. Armstrong, Jasmine Perdeh, Bjorn Bauer, Jeffrey Talbert, V. K. Cody Bumgardner

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.09057 2024-01-30 cs.CV 79%

CrossVideo: Self-supervised Cross-modal Contrastive Learning for Point Cloud Video Understanding

Yunze Liu, Changxi Chen, Zifan Wang, Li Yi

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

Journal ref ICRA2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.11649 2024-01-23 cs.CV 79%

M2-CLIP: A Multimodal, Multi-task Adapting Framework for Video Action Recognition

Mengmeng Wang, Jiazheng Xing, Boyuan Jiang, Jun Chen, Jianbiao Mei, Xingxing Zuo, Guang Dai, Jingdong Wang, Yong Liu

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Journal ref AAAI2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.01061 2024-01-22 cs.CV 79%

Rethinking Cross-modal Interaction from a Top-down Perspective for Referring Video Object Segmentation

Chen Liang, Yu Wu, Tianfei Zhou, Wenguan Wang, Zongxin Yang, Yunchao Wei, Yi Yang

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

Comments Champion solution in YouTube-VOS 2021 Track 3. Extended version published in https://ieeexplore.ieee.org/abstract/document/10083244

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.08345 2024-01-17 cs.CV 79%

Multi-view Distillation based on Multi-modal Fusion for Few-shot Action Recognition(CLIP-$\mathrm{M^2}$DF)

Fei Guo, YiKang Wang, Han Qi, WenPing Jin, Li Zhu

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.08043 2024-01-17 cs.RO cs.CV 79%

Cross-Modal Semi-Dense 6-DoF Tracking of an Event Camera in Challenging Conditions

Yi-Fan Zuo, Wanting Xu, Xia Wang, Yifu Wang, Laurent Kneip

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

Comments accepted by IEEE Transactions on Robotics (T-RO). arXiv admin note: text overlap with arXiv:2202.02556

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.07218 2024-01-17 cs.CV 79%

Self-supervised Event-based Monocular Depth Estimation using Cross-modal Consistency

Junyu Zhu, Lina Liu, Bofeng Jiang, Feng Wen, Hongbo Zhang, Wanlong Li, Yong Liu

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

Comments Accepted by IROS2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.06942 2024-01-05 cs.CV 79%

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Yi Wang, Yinan He, Yizhuo Li, Kunchang Li, Jiashuo Yu, Xin Ma, Xinhao Li, Guo Chen, Xinyuan Chen, Yaohui Wang, Conghui He, Ping Luo, Ziwei Liu, Yali Wang, Limin Wang, Yu Qiao

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments Data and Code: https://github.com/OpenGVLab/InternVideo/tree/main/Data/InternVid

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.00988 2024-01-03 cs.CV 79%

Holistic Autonomous Driving Understanding by Bird's-Eye-View Injected Multi-Modal Large Models

Xinpeng Ding, Jinahua Han, Hang Xu, Xiaodan Liang, Wei Zhang, Xiaomeng Li

专题命中 视频多模态 :multi-modal(title);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.15646 2023-12-29 cs.CY cs.AI econ.GN q-fin.EC 79%

A graph-based multimodal framework to predict gentrification

Javad Eshtiyagh, Baotong Zhang, Yujing Sun, Linhui Wu, Zhao Wang

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

Journal ref International Conference on Urban Informatics 2023 - Best Paper Award 3rd Place

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.13633 2023-12-22 cs.CV 79%

Multi-Modal Domain Adaptation Across Video Scenes for Temporal Video Grounding

Haifeng Huang, Yang Zhao, Zehan Wang, Yan Xia, Zhou Zhao

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.11935 2023-12-20 cs.AI 79%

Parameterized Decision-making with Multi-modal Perception for Autonomous Driving

Yuyang Xia, Shuncheng Liu, Quanlin Yu, Liwei Deng, You Zhang, Han Su, Kai Zheng

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.AI

Comments IEEE International Conference on Data Engineering (ICDE2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.03056 2023-12-20 cs.CV 79%

MOISST: Multimodal Optimization of Implicit Scene for SpatioTemporal calibration

Quentin Herau, Nathan Piasco, Moussab Bennehar, Luis Roldão, Dzmitry Tsishkou, Cyrille Migniot, Pascal Vasseur, Cédric Demonceaux

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments Accepted at IROS2023 Project site: https://qherau.github.io/MOISST/

详情

展开后加载摘要…

URL PDF HTML 收藏