arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6918 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6918 篇

2208.00394 2023-07-07 cs.CV cs.AI 81%

STrajNet: Multi-modal Hierarchical Transformer for Occupancy Flow Field Prediction in Autonomous Driving

Haochen Liu, Zhiyu Huang, Chen Lv

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.12776 2023-07-06 cs.CV cs.AI 81%

SFusion: Self-attention based N-to-One Multimodal Fusion Block

Zecheng Liu, Jia Wei, Rui Li, Jianlong Zhou

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments This paper has been accepted by MICCAI 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.16645 2023-06-30 cs.CV cs.MM 81%

Deep Equilibrium Multimodal Fusion

Jinhong Ni, Yalong Bai, Wei Zhang, Ting Yao, Tao Mei

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.07935 2023-06-14 cs.CL cs.AI cs.LG 81%

Multi-modal Representation Learning for Social Post Location Inference

Ruiting Dai, Jiayi Luo, Xucheng Luo, Lisi Mo, Wanlun Ma, Fan Zhou

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CL、cs.AI

Comments 6 pages, 2023 International Conference on Communications

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.04790 2023-06-14 cs.CV cs.CL 81%

MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

Tao Gong, Chengqi Lyu, Shilong Zhang, Yudong Wang, Miao Zheng, Qian Zhao, Kuikun Liu, Wenwei Zhang, Ping Luo, Kai Chen

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments 10 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.04387 2023-06-09 cs.CV cs.CL 81%

M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Lei Li, Yuwei Yin, Shicheng Li, Liang Chen, Peiyi Wang, Shuhuai Ren, Mukai Li, Yazheng Yang, Jingjing Xu, Xu Sun, Lingpeng Kong, Qi Liu

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV、cs.CL

Comments Fix dataset url: https://huggingface.co/datasets/MMInstruction/M3IT Project: https://m3-it.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.00864 2023-06-02 cs.CV cs.CL cs.LG 81%

A Transformer-based representation-learning model with unified processing of multimodal input for clinical diagnostics

Hong-Yu Zhou, Yizhou Yu, Chengdi Wang, Shu Zhang, Yuanxu Gao, Jia Pan, Jun Shao, Guangming Lu, Kang Zhang, Weimin Li

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Accepted by Nature Biomedical Engineering

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.12248 2023-05-23 cs.CL cs.CV 81%

Brain encoding models based on multimodal transformers can transfer across language and vision

Jerry Tang, Meng Du, Vy A. Vo, Vasudev Lal, Alexander G. Huth

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.11303 2023-05-11 cs.CL cs.CV 81%

Modeling Paragraph-Level Vision-Language Semantic Alignment for Multi-Modal Summarization

Chenhao Cui, Xinnian Liang, Shuangzhi Wu, Zhoujun Li

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.11097 2023-05-01 cs.AI cs.CV 81%

A Multi-Modal Neural Geometric Solver with Textual Clauses Parsed from Diagram

Ming-Liang Zhang, Fei Yin, Cheng-Lin Liu

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted to IJCAI 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.13725 2023-04-28 eess.IV cs.AI cs.CV cs.LG 81%

Prediction of brain tumor recurrence location based on multi-modal fusion and nonlinear correlation learning

Tongxue Zhou, Alexandra Noeuveglise, Romain Modzelewski, Fethi Ghazouani, Sébastien Thureau, Maxime Fontanilles, Su Ruan

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments 23 pages, 4 figures

Journal ref Computerized Medical Imaging and Graphics, 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.00827 2023-04-04 cs.CV cs.MM 81%

Multi-modal Fake News Detection on Social Media via Multi-grained Information Fusion

Yangming Zhou, Yuzhou Yang, Qichao Ying, Zhenxing Qian, Xinpeng Zhang

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV、cs.MM

Comments Accepted by ICMR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.17611 2023-04-03 cs.HC cs.AI cs.SD eess.AS 81%

Transformer-based Self-supervised Multimodal Representation Learning for Wearable Emotion Recognition

Yujin Wu, Mohamed Daoudi, Ali Amad

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI、eess.AS

Comments Accepted IEEE Transactions On Affective Computing

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.14708 2023-03-28 cs.CV cs.MM 81%

Exploring Multimodal Sentiment Analysis via CBAM Attention and Double-layer BiLSTM Architecture

Huiru Wang, Xiuhong Li, Zenyu Ren, Dan Yang, chunming Ma

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.08774 2023-03-14 cs.AI cs.MM 81%

Vision, Deduction and Alignment: An Empirical Study on Multi-modal Knowledge Graph Alignment

Yangning Li, Jiaoyan Chen, Yinghui Li, Yuejia Xiang, Xi Chen, Hai-Tao Zheng

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.AI、cs.MM

Comments Accepted by ICASSP2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.13661 2023-02-28 cs.CL cs.SD eess.AS 81%

Using Auxiliary Tasks In Multimodal Fusion Of Wav2vec 2.0 And BERT For Multimodal Emotion Recognition

Dekai Sun, Yancheng He, Jiqing Han

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.15588 2023-01-30 cs.LG cs.AI cs.CV 81%

Deep Multi-modal Fusion of Image and Non-image Data in Disease Diagnosis and Prognosis: A Review

Can Cui, Haichun Yang, Yaohong Wang, Shilin Zhao, Zuhayr Asad, Lori A. Coburn, Keith T. Wilson, Bennett A. Landman, Yuankai Huo

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.10419 2023-01-23 cs.LG cs.AI cs.CV cs.RO 81%

Learning Sequential Latent Variable Models from Multimodal Time Series Data

Oliver Limoyo, Trevor Ablett, Jonathan Kelly

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments In: Petrovic, I., Menegatti, E., Marković, I. (eds) Intelligent Autonomous Systems 17. IAS 2022. Lecture Notes in Networks and Systems, vol 577. Springer, Cham

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.00676 2023-01-03 cs.LG cs.AI cs.CL 81%

Multimodal Sequential Generative Models for Semi-Supervised Language Instruction Following

Kei Akuzawa, Yusuke Iwasawa, Yutaka Matsuo

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.15837 2022-11-30 cs.LG cs.AI cs.CV cs.GT 81%

Survey on Self-Supervised Multimodal Representation Learning and Foundation Models

Sushil Thapa

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.11925 2022-11-23 cs.CV cs.AI 81%

Multimodal Data Augmentation for Visual-Infrared Person ReID with Corrupted Data

Arthur Josi, Mahdi Alehdaghi, Rafael M. O. Cruz, Eric Granger

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 8 pages of main content, 2 pages of references, 2 pages of supplementary material, 3 figures, WACV 2023 RWS workshop,

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.08821 2022-11-02 cs.CL cs.MM 81%

MoSE: Modality Split and Ensemble for Multimodal Knowledge Graph Completion

Yu Zhao, Xiangrui Cai, Yike Wu, Haiwei Zhang, Ying Zhang, Guoqing Zhao, Ning Jiang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL、cs.MM

Comments Accepted by EMNLP 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.09946 2022-11-01 cs.MM cs.AI cs.LG 81%

MMGA: Multimodal Learning with Graph Alignment

Xuan Yang, Quanjin Tao, Xiao Feng, Donghong Cai, Xiang Ren, Yang Yang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI、cs.MM

Comments Please contact xuany@zju.edu.cn for the dataset

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.10972 2022-10-25 cs.MM cs.CV 81%

A Multimodal Sensor Fusion Framework Robust to Missing Modalities for Person Recognition

Vijay John, Yasutomo Kawanishi

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.MM

Comments Accepted for ACM Multimedia Asia, 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.11024 2022-10-21 cs.LG cs.AI cs.CV 81%

A survey on Self Supervised learning approaches for improving Multimodal representation learning

Naman Goyal

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.05790 2022-10-13 cs.LG cs.CL cs.CV 81%

Transfer Learning with Joint Fine-Tuning for Multimodal Sentiment Analysis

Guilherme Lourenço de Toledo, Ricardo Marcondes Marcacini

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Talk: https://icml.cc/Conferences/2022/ScheduleMultitrack?event=13483#collapse20429

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.04510 2022-10-11 cs.CV cs.AI 81%

Multi-Modal Fusion Transformer for Visual Question Answering in Remote Sensing

Tim Siebert, Kai Norman Clasen, Mahdyar Ravanbakhsh, Begüm Demir

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted in SPIE Remote Sensing (ESI22R)

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.03289 2022-10-10 cs.CV cs.AI cs.LG 81%

Scalable Self-Supervised Representation Learning from Spatiotemporal Motion Trajectories for Multimodal Computer Vision

Swetava Ganguli, C. V. Krishnakumar Iyer, Vipul Pandey

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Extended abstract accepted for presentation at BayLearn 2022. 3 pages, 2 figures, 1 table. Abstract based on IEEE MDM 2022 research track paper: arXiv:2110.12521

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.11893 2022-08-26 cs.CL cs.AI 81%

Cross-Modality Gated Attention Fusion for Multimodal Sentiment Analysis

Ming Jiang, Shaoxiong Ji

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.12521 2022-07-19 cs.CV cs.AI cs.LG 81%

Reachability Embeddings: Scalable Self-Supervised Representation Learning from Mobility Trajectories for Multimodal Geospatial Computer Vision

Swetava Ganguli, C. V. Krishnakumar Iyer, Vipul Pandey

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Extended version of the accepted research track paper at the 23rd IEEE International Conference on Mobile Data Management (MDM), 2022, Paphos, Cyprus. 12 pages, 6 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏