arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6918 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6918 篇

2304.11381 2023-04-25 cs.CV 79%

Incomplete Multimodal Learning for Remote Sensing Data Fusion

Yuxing Chen, Maofan Zhao, Lorenzo Bruzzone

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.09694 2023-04-20 cs.CV 79%

CrossFusion: Interleaving Cross-modal Complementation for Noise-resistant 3D Object Detection

Yang Yang, Weijie Ma, Hao Chen, Linlin Ou, Xinyi Yu

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.07675 2023-04-18 cs.CV 79%

Multimodal Representation Learning of Cardiovascular Magnetic Resonance Imaging

Jielin Qiu, Peide Huang, Makiya Nakashima, Jaehyun Lee, Jiacheng Zhu, Wilson Tang, Pohao Chen, Christopher Nguyen, Byung-Hak Kim, Debbie Kwon, Douglas Weber, Ding Zhao, David Chen

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments 24 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.07147 2023-04-17 eess.IV cs.CV cs.LG 79%

Cross Attention Transformers for Multi-modal Unsupervised Whole-Body PET Anomaly Detection

Ashay Patel, Petru-Danial Tudiosu, Walter H. L. Pinaya, Gary Cook, Vicky Goh, Sebastien Ourselin, M. Jorge Cardoso

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted for publication at the Journal of Machine Learning for Biomedical Imaging (MELBA) https://melba-journal.org/2023:006

Journal ref Machine.Learning.for.Biomedical.Imaging. 2 (2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.09164 2023-04-17 cs.CV 79%

Multimodal Feature Extraction and Fusion for Emotional Reaction Intensity Estimation and Expression Classification in Videos with Transformers

Jia Li, Yin Chen, Xuesong Zhang, Jiantao Nie, Ziqiang Li, Yangchen Yu, Yan Zhang, Richang Hong, Meng Wang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Solutions of HFUT-CVers Team at the 5th ABAW Competition (CVPR 2023 workshop)

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.06485 2023-04-14 eess.SP cs.AI cs.LG 79%

CoRe-Sleep: A Multimodal Fusion Framework for Time Series Robust to Imperfect Modalities

Konstantinos Kontras, Christos Chatzichristos, Huy Phan, Johan Suykens, Maarten De Vos

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments 10 pages, 4 figures, 2 tables, journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.02853 2023-04-07 cs.CV 79%

Learning Instance-Level Representation for Large-Scale Multi-Modal Pretraining in E-commerce

Yang Jin, Yongzhi Li, Zehuan Yuan, Yadong Mu

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments 16 pages, 10 figures, accepted by CVPR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.01563 2023-04-05 cs.CL 79%

Attribute-Consistent Knowledge Graph Representation Learning for Multi-Modal Entity Alignment

Qian Li, Shu Guo, Yangyifei Luo, Cheng Ji, Lihong Wang, Jiawei Sheng, Jianxin Li

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.13810 2023-03-27 cs.CV eess.IV 79%

Evidence-aware multi-modal data fusion and its application to total knee replacement prediction

Xinwen Liu, Jing Wang, S. Kevin Zhou, Craig Engstrom, Shekhar S. Chandra

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.13802 2023-03-27 cs.CV 79%

Decoupled Multimodal Distilling for Emotion Recognition

Yong Li, Yuanzhi Wang, Zhen Cui

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments To appear at CVPR 2023, selected as a hightlight, 10% of accepted papers, 2.5% of submissions

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.13101 2023-03-24 cs.CV 79%

MMFormer: Multimodal Transformer Using Multiscale Self-Attention for Remote Sensing Image Classification

Bo Zhang, Zuheng Ming, Wei Feng, Yaqian Liu, Liang He, Kaixing Zhao

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.01392 2023-03-24 cs.CV 79%

Multi-modal Gated Mixture of Local-to-Global Experts for Dynamic Image Fusion

Yiming Sun, Bing Cao, Pengfei Zhu, Qinghua Hu

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.10816 2023-03-21 cs.AI cs.IR 79%

IMF: Interactive Multimodal Fusion Model for Link Prediction

Xinhang Li, Xiangyu Zhao, Jiaxing Xu, Yong Zhang, Chunxiao Xing

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments 9 pages, 7 figures, 4 tables, WWW'2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.07126 2023-03-14 eess.IV cs.CV 79%

Mirror U-Net: Marrying Multimodal Fission with Multi-task Learning for Semantic Segmentation in Medical Imaging

Zdravko Marinov, Simon Reiß, David Kersting, Jens Kleesiek, Rainer Stiefelhagen

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments 8 pages; 8 figures; 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.05302 2023-03-10 eess.IV cs.CV 79%

M3AE: Multimodal Representation Learning for Brain Tumor Segmentation with Missing Modalities

Hong Liu, Dong Wei, Donghuan Lu, Jinghan Sun, Liansheng Wang, Yefeng Zheng

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Journal ref AAAI 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.12735 2023-03-08 cs.CV 79%

Multi-Modal 3D Object Detection in Autonomous Driving: a Survey

Yingjie Wang, Qiuyu Mao, Hanqi Zhu, Jiajun Deng, Yu Zhang, Jianmin Ji, Houqiang Li, Yanyong Zhang

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted by International Journal of Computer Vision (IJCV)

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.11393 2023-03-03 cs.CV 79%

TFormer: A throughout fusion transformer for multi-modal skin lesion diagnosis

Yilan Zhang, Fengying Xie, Jianqi Chen

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments 16 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.03028 2023-03-03 eess.IV cs.CV 79%

Multimodal Brain Disease Classification with Functional Interaction Learning from Single fMRI Volume

Wei Dai, Ziyao Zhang, Lixia Tian, Shengyuan Yu, Shuhui Wang, Zhao Dong, Hairong Zheng

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.00403 2023-03-02 cs.CV cs.LG 79%

Can representation learning for multimodal image registration be improved by supervision of intermediate layers?

Elisabeth Wetzer, Joakim Lindblad, Nataša Sladoje

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments 15 Pages + 9 Pages Appendix, 10 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.09318 2023-02-21 cs.LG cs.AI 79%

Effective Multimodal Reinforcement Learning with Modality Alignment and Importance Enhancement

Jinming Ma, Feng Wu, Yingfeng Chen, Xianpeng Ji, Yu Ding

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments 10 pages, 12 figures, This article is an extended version of the Extended Abstract accepted by the International Conference on Autonomous Agents and Multi-Agent Systems (AAMAS-2022)

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.08670 2023-02-20 cs.CV cs.IR 79%

Cascaded information enhancement and cross-modal attention feature fusion for multispectral pedestrian detection

Yang Yang, Kaixiong Xu, Kaizheng Wang

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.08326 2023-02-17 cs.CL 79%

NUAA-QMUL-AIIT at Memotion 3: Multi-modal Fusion with Squeeze-and-Excitation for Internet Meme Emotion Analysis

Xiaoyu Guo, Jing Ma, Arkaitz Zubiaga

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.05744 2023-02-14 cs.CV 79%

Rethinking Vision Transformer and Masked Autoencoder in Multimodal Face Anti-Spoofing

Zitong Yu, Rizhao Cai, Yawen Cui, Xin Liu, Yongjian Hu, Alex Kot

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.12097 2023-02-10 cs.IR cs.MM 79%

Enhancing Dyadic Relations with Homogeneous Graphs for Multimodal Recommendation

Hongyu Zhou, Xin Zhou, Lingzi Zhang, Zhiqi Shen

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.MM

Comments 17 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.02277 2023-02-08 cs.CV 79%

R2FD2: Fast and Robust Matching of Multimodal Remote Sensing Image via Repeatable Feature Detector and Rotation-invariant Feature Descriptor

Bai Zhu, Chao Yang, Jinkun Dai, Jianwei Fan, Yuanxin Ye

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments 33 pages, 15 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.00941 2023-02-08 cs.CV eess.IV eess.SP 79%

Unsupervised Multimodal Change Detection Based on Structural Relationship Graph Representation Learning

Hongruixuan Chen, Naoto Yokoya, Chen Wu, Bo Du

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.08320 2023-02-03 cs.CV 79%

Autoencoders as Cross-Modal Teachers: Can Pretrained 2D Image Transformers Help 3D Representation Learning?

Runpei Dong, Zekun Qi, Linfeng Zhang, Junbo Zhang, Jianjian Sun, Zheng Ge, Li Yi, Kaisheng Ma

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

Comments Accepted at ICLR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.07502 2023-01-24 cs.LG cs.CV 79%

Multimodal Side-Tuning for Document Classification

Stefano Pio Zingaro, Giuseppe Lisanti, Maurizio Gabbrielli

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments 2020 25th International Conference on Pattern Recognition (ICPR)

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.03033 2023-01-10 cs.CV 79%

RGB-T Multi-Modal Crowd Counting Based on Transformer

Zhengyi Liu, Wei Wu, Yacheng Tan, Guanghui Zhang

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Journal ref BMVC2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.00933 2023-01-02 cs.CV 79%

Deep Multimodal Fusion for Generalizable Person Re-identification

Suncheng Xiang, Hao Chen, Wei Ran, Zefang Yu, Ting Liu, Dahong Qian, Yuzhuo Fu

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏