arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4749 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4749 篇

2312.07378 2023-12-13 cs.CV 79%

X4D-SceneFormer: Enhanced Scene Understanding on 4D Point Cloud Videos through Cross-modal Knowledge Transfer

Linglin Jing, Ying Xue, Xu Yan, Chaoda Zheng, Dong Wang, Ruimao Zhang, Zhigang Wang, Hui Fang, Bin Zhao, Zhen Li

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.19062 2023-11-28 cs.RO cs.AI 79%

A multi-modal table tennis robot system

Andreas Ziegler, Thomas Gossard, Karl Vetter, Jonas Tebbe, Andreas Zell

专题命中 视频多模态 :multi-modal(title);multimodal(abstract);分类 cs.AI

Comments Accepted for RoboLetics: Workshop on Robot Learning in Athletics @CoRL 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.12344 2023-11-22 cs.CV 79%

Modality Mixer Exploiting Complementary Information for Multi-modal Action Recognition

Sumin Lee, Sangmin Woo, Muhammad Adi Nugroho, Changick Kim

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments arXiv admin note: substantial text overlap with arXiv:2208.11314

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.07608 2023-11-15 cs.LG cs.AI 79%

MuST: Multimodal Spatiotemporal Graph-Transformer for Hospital Readmission Prediction

Yan Miao, Lequan Yu

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.20201 2023-11-01 cs.CL 79%

Video-Helpful Multimodal Machine Translation

Yihang Li, Shuichiro Shimizu, Chenhui Chu, Sadao Kurohashi, Wei Li

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL

Comments Accepted by EMNLP 2023 Main Conference (long paper)

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.16590 2023-10-26 cs.CV 79%

$\mathbb{VD}$-$\mathbb{GR}$: Boosting $\mathbb{V}$isual $\mathbb{D}$ialog with Cascaded Spatial-Temporal Multi-Modal $\mathbb{GR}$aphs

Adnen Abdessaied, Lei Shi, Andreas Bulling

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments WACV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.15670 2023-10-25 cs.CV 79%

Leveraging Vision-Centric Multi-Modal Expertise for 3D Object Detection

Linyan Huang, Zhiqi Li, Chonghao Sima, Wenhai Wang, Jingdong Wang, Yu Qiao, Hongyang Li

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted by NeurIPS 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.11169 2023-10-18 cs.LG cs.AI 79%

MST-GAT: A Multimodal Spatial-Temporal Graph Attention Network for Time Series Anomaly Detection

Chaoyue Ding, Shiliang Sun, Jing Zhao

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

Comments Information Fusion 2023 accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.09222 2023-10-13 cs.CV cs.LG 79%

MMTSA: Multimodal Temporal Segment Attention Network for Efficient Human Activity Recognition

Ziqi Gao, Yuntao Wang, Jianguo Chen, Junliang Xing, Shwetak Patel, Xin Liu, Yuanchun Shi

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.04991 2023-10-12 cs.CV 79%

Video-Teller: Enhancing Cross-Modal Generation with Fusion and Decoupling

Haogeng Liu, Qihang Fan, Tingkai Liu, Linjie Yang, Yunzhe Tao, Huaibo Huang, Ran He, Hongxia Yang

专题命中 视频多模态 :cross-modal(title);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.09067 2023-09-20 cs.CV 79%

MMST-ViT: Climate Change-aware Crop Yield Prediction via Multi-Modal Spatial-Temporal Vision Transformer

Fudong Lin, Summer Crawford, Kaleb Guillot, Yihe Zhang, Yan Chen, Xu Yuan, Li Chen, Shelby Williams, Robert Minvielle, Xiangming Xiao, Drew Gholson, Nicolas Ashwell, Tri Setiyono, Brenda Tubana, Lu Peng, Magdy Bayoumi, Nian-Feng Tzeng

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Journal ref ICCV 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.02320 2023-09-06 physics.geo-ph cs.AI cs.LG 79%

SeisCLIP: A seismology foundation model pre-trained by multi-modal data for multi-purpose seismic feature extraction

Xu Si, Xinming Wu, Hanlin Sheng, Jun Zhu, Zefeng Li

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.AI

Comments 27 pages, 9 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.04129 2023-09-06 cs.CV 79%

Cross-modal Orthogonal High-rank Augmentation for RGB-Event Transformer-trackers

Zhiyu Zhu, Junhui Hou, Dapeng Oliver Wu

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

Comments accepted by ICCV

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.14383 2023-08-29 cs.CV 79%

Multi-Modal Neural Radiance Field for Monocular Dense SLAM with a Light-Weight ToF Sensor

Xinyang Liu, Yijin Li, Yanbin Teng, Hujun Bao, Guofeng Zhang, Yinda Zhang, Zhaopeng Cui

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted to ICCV 2023 (Oral). Project Page: https://zju3dv.github.io/tof_slam/

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.07205 2023-08-29 cs.CV 79%

Multimodal Motion Conditioned Diffusion Model for Skeleton-based Video Anomaly Detection

Alessandro Flaborea, Luca Collorone, Guido D'Amely, Stefano D'Arrigo, Bardh Prenkaj, Fabio Galasso

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments Accepted at ICCV2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.11185 2023-08-23 cs.CV 79%

MEGA: Multimodal Alignment Aggregation and Distillation For Cinematic Video Segmentation

Najmeh Sadoughi, Xinyu Li, Avijit Vajpayee, David Fan, Bing Shuai, Hector Santos-Villalobos, Vimal Bhat, Rohith MV

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments ICCV 2023 accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.05638 2023-08-21 cs.CV 79%

Cross-Modal Learning with 3D Deformable Attention for Action Recognition

Sangwon Kim, Dasom Ahn, Byoung Chul Ko

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

Comments Accepted by ICCV2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.02459 2023-08-07 cs.RO cs.AI cs.LG 79%

Nonprehensile Planar Manipulation through Reinforcement Learning with Multimodal Categorical Exploration

Juan Del Aguila Ferrandis, João Moura, Sethu Vijayakumar

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.03906 2023-07-31 cs.CV 79%

Cross-modal Manifold Cutmix for Self-supervised Video Representation Learning

Srijan Das, Michael S. Ryoo

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

Comments Accepted at MVA 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.10346 2023-07-21 cs.HC cs.MM 79%

Estudio de la Experiencia de Usuario mediante un Sistema de Dashboards de Análisis de Aprendizaje Multimodal

Álvaro Becerra, Roberto Daza, Ruth Cobos, Aythami Morales, Julian Fierrez

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.MM

Comments Accepted in "XXIII CONGRESO INTERNACIONAL DE INTERACCIÓN PERSONA-ORDENADOR 2023". Article in Spanish language. The abstract in English and Spanish. There is an extended abstract of 2 pages in English

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.07807 2023-07-18 eess.IV cs.CV 79%

MUVF-YOLOX: A Multi-modal Ultrasound Video Fusion Network for Renal Tumor Diagnosis

Junyu Li, Han Huang, Dong Ni, Wufeng Xue, Dongmei Zhu, Jun Cheng

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments MICCAI 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.01047 2023-07-04 cs.CV 79%

Cross-modal Place Recognition in Image Databases using Event-based Sensors

Xiang Ji, Jiaxin Wei, Yifu Wang, Huiliang Shang, Laurent Kneip

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.00226 2023-07-04 cs.CV cs.LG 79%

S-Omninet: Structured Data Enhanced Universal Multimodal Learning Architecture

Ye Xue, Diego Klabjan, Jean Utke

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.14392 2023-06-27 cs.CV 79%

ContentCTR: Frame-level Live Streaming Click-Through Rate Prediction with Multimodal Transformer

Jiaxin Deng, Dong Shen, Shiyao Wang, Xiangyu Wu, Fan Yang, Guorui Zhou, Gaofeng Meng

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.12156 2023-06-07 cs.LG cs.AI 79%

Improving Medical Predictions by Irregular Multimodal Electronic Health Records Modeling

Xinlu Zhang, Shiyang Li, Zhiyu Chen, Xifeng Yan, Linda Petzold

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.02845 2023-06-06 cs.AI 79%

Interpretable Multimodal Emotion Recognition using Facial Features and Physiological Signals

Puneet Kumar, Xiaobai Li

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

Comments Accepted for Oral Presentation in DAI 2023 (https://rbcdsai.iitm.ac.in/DAI-2023/program.html)

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.00640 2023-06-02 cs.CV eess.IV 79%

Multi-Modal Deep Learning for Multi-Temporal Urban Mapping With a Partly Missing Optical Modality

Sebastian Hafner, Yifang Ban

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments 4 pages, 2 figures, accepted for publication in the IGARSS 2023 Proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.19624 2023-06-01 cs.CV 79%

A Multi-Modal Transformer Network for Action Detection

Matthew Korban, Scott T. Acton, Peter Youngs

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Journal ref Pattern Recognition 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.08352 2023-05-31 cs.CV 79%

MHSCNet: A Multimodal Hierarchical Shot-aware Convolutional Network for Video Summarization

Wujiang Xu, Runzhong Wang, Xiaobo Guo, Shaoshuai Li, Qiongxu Ma, Yunan Zhao, Sheng Guo, Zhenfeng Zhu, Junchi Yan

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.12448 2023-05-26 cs.CV 79%

CMD: Self-supervised 3D Action Representation Learning with Cross-modal Mutual Distillation

Yunyao Mao, Wengang Zhou, Zhenbo Lu, Jiajun Deng, Houqiang Li

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

Comments To appear in ECCV 2022 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏