arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6918 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6918 篇

2310.02663 2023-10-05 cs.CV 79%

MedPrompt: Cross-Modal Prompting for Multi-Task Medical Image Translation

Xuhang Chen, Chi-Man Pun, Shuqiang Wang

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.01912 2023-10-04 eess.IV cs.CV cs.LG 79%

Improved Automatic Diabetic Retinopathy Severity Classification Using Deep Multimodal Fusion of UWF-CFP and OCTA Images

Mostafa El Habib Daho, Yihao Li, Rachid Zeghlache, Yapo Cedric Atse, Hugo Le Boité, Sophie Bonnin, Deborah Cosette, Pierre Deman, Laurent Borderie, Capucine Lepicard, Ramin Tadayoni, Béatrice Cochener, Pierre-Henri Conze, Mathieu Lamard, Gwenolé Quellec

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Accepted preprint for presentation at MICCAI-OMIA 20023, Vancouver, Canada

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.06612 2023-09-29 cs.LG cs.CV 79%

Harmonic-NAS: Hardware-Aware Multimodal Neural Architecture Search on Resource-constrained Devices

Mohamed Imed Eddine Ghebriout, Halima Bouzidi, Smail Niar, Hamza Ouarnoughi

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Accepted to the 15th Asian Conference on Machine Learning (ACML 2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.15529 2023-09-28 eess.IV cs.CV cs.LG 79%

Missing-modality Enabled Multi-modal Fusion Architecture for Medical Data

Muyu Wang, Shiyu Fan, Yichen Li, Hui Chen

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.13650 2023-09-26 eess.AS cs.SD 79%

Cross-modal Alignment with Optimal Transport for CTC-based ASR

Xugang Lu, Peng Shen, Yu Tsao, Hisashi Kawai

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 eess.AS

Comments Accepted to IEEE ASRU 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.11811 2023-09-22 eess.SP cs.AI 79%

Multimodal Transformers for Wireless Communications: A Case Study in Beam Prediction

Yu Tian, Qiyang Zhao, Zine el abidine Kherroubi, Fouzi Boukhalfa, Kebin Wu, Faouzi Bader

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.09593 2023-09-19 cs.CV cs.IT cs.RO math.IT 79%

Mutual Information-calibrated Conformal Feature Fusion for Uncertainty-Aware Multimodal 3D Object Detection at the Edge

Alex C. Stutts, Danilo Erricolo, Sathya Ravi, Theja Tulabandhula, Amit Ranjan Trivedi

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.05787 2023-09-13 cs.AI cs.HC cs.LG 79%

Adaptive User-centered Neuro-symbolic Learning for Multimodal Interaction with Autonomous Systems

Amr Gomaa, Michael Feld

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments AI&HCI Workshop accepted paper at ICML2023 and accepted at ICMI2023 Blue Sky Papers. arXiv admin note: text overlap with arXiv:2211.03539

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.05608 2023-09-12 cs.CL cs.CE 79%

Incorporating Pre-trained Model Prompting in Multimodal Stock Volume Movement Prediction

Ruibo Chen, Zhiyuan Zhang, Yi Liu, Ruihan Bao, Keiko Harimoto, Xu Sun

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

Comments 9 pages, 3 figures, 7 tables. Accepted by 2023 KDD Workshop on Machine Learning in Finance

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.07901 2023-09-12 cs.LG cs.CV 79%

Auxiliary Cross-Modal Representation Learning with Triplet Loss Functions for Online Handwriting Recognition

Felix Ott, David Rügamer, Lucas Heublein, Bernd Bischl, Christopher Mutschler

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

Journal ref IEEE Access, volume 11, pages 94148-94172, August 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.04062 2023-09-11 cs.LG cs.AI physics.chem-ph 79%

3D Denoisers are Good 2D Teachers: Molecular Pretraining via Denoising and Cross-Modal Distillation

Sungjun Cho, Dae-Woong Jeong, Sung Moon Ko, Jinwoo Kim, Sehui Han, Seunghoon Hong, Honglak Lee, Moontae Lee

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.AI

Comments 16 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.00601 2023-09-08 cs.CV 79%

Multimodal Industrial Anomaly Detection via Hybrid Fusion

Yue Wang, Jinlong Peng, Jiangning Zhang, Ran Yi, Yabiao Wang, Chengjie Wang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Accepted by CVPR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.02702 2023-09-07 cs.CV 79%

Gene-induced Multimodal Pre-training for Image-omic Classification

Ting Jin, Xingran Xie, Renjie Wan, Qingli Li, Yan Wang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.01169 2023-09-06 cs.LG cs.AI 79%

End-to-End Learning on Multimodal Knowledge Graphs

W. X. Wilcke, P. Bloem, V. de Boer, R. H. van t Veer

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments Under submission. arXiv admin note: substantial text overlap with arXiv:2003.12383

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.09941 2023-09-04 cs.CV 79%

A Robust and Interpretable Deep Learning Framework for Multi-modal Registration via Keypoints

Alan Q. Wang, Evan M. Yu, Adrian V. Dalca, Mert R. Sabuncu

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted to Medical Image Analysis 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.14505 2023-09-01 eess.IV cs.CV cs.LG 79%

Transformer-based interpretable multi-modal data fusion for skin lesion classification

Theodor Cheslerean-Boghiu, Melia-Evelina Fleischmann, Theresa Willem, Tobias Lasser

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments Submitted to IEEE JBHI in July 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.15846 2023-08-31 cs.CV 79%

Exploring Multi-Modal Contextual Knowledge for Open-Vocabulary Object Detection

Yifan Xu, Mengdan Zhang, Xiaoshan Yang, Changsheng Xu

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.13883 2023-08-29 eess.IV cs.CV 79%

ReFuSeg: Regularized Multi-Modal Fusion for Precise Brain Tumour Segmentation

Aditya Kasliwal, Sankarshanaa Sagaram, Laven Srivastava, Pratinav Seth, Adil Khan

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted at 9th edition of the Brain Lesion (BrainLes) workshop, MICCAI 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.11217 2023-08-25 cs.LG cs.AI 79%

Federated Learning in Big Model Era: Domain-Specific Multimodal Large Models

Zengxiang Li, Zhaoxiang Hou, Hui Liu, Ying Wang, Tongzhi Li, Longfei Xie, Chao Shi, Chengyi Yang, Weishan Zhang, Zelei Liu, Liang Xu

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.10741 2023-08-22 cs.LG cs.AI cs.CR 79%

On the Adversarial Robustness of Multi-Modal Foundation Models

Christian Schlarmann, Matthias Hein

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.AI

Comments ICCV AROW 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.07971 2023-08-17 cs.CL cs.LG 79%

MultiSChuBERT: Effective Multimodal Fusion for Scholarly Document Quality Prediction

Gideon Maillette de Buy Wenniger, Thomas van Dongen, Lambert Schomaker

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.06794 2023-08-16 cs.CV 79%

MMF-Track: Multi-modal Multi-level Fusion for 3D Single Object Tracking

Zhiheng Li, Yubo Cui, Yu Lin, Zheng Fang

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments 11 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.06530 2023-08-15 cs.CV 79%

BEV-DG: Cross-Modal Learning under Bird's-Eye View for Domain Generalization of 3D Semantic Segmentation

Miaoyu Li, Yachao Zhang, Xu MA, Yanyun Qu, Yun Fu

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

Comments Accepted by ICCV 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.04502 2023-08-15 cs.CL 79%

Revisiting Disentanglement and Fusion on Modality and Context in Conversational Multimodal Emotion Recognition

Bobo Li, Hao Fei, Lizi Liao, Yu Zhao, Chong Teng, Tat-Seng Chua, Donghong Ji, Fei Li

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

Comments Accepted by ACM MM 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.03113 2023-08-08 cs.IR cs.MM 79%

Semantic-Guided Feature Distillation for Multimodal Recommendation

Fan Liu, Huilin Chen, Zhiyong Cheng, Liqiang Nie, Mohan Kankanhalli

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.MM

Comments ACM Multimedia 2023 Accepted

Journal ref In Proceedings of the 31st ACM International Conference on Multimedia (MM '23), 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.01430 2023-08-04 cs.CL 79%

FinVis-GPT: A Multimodal Large Language Model for Financial Chart Analysis

Ziao Wang, Yuhang Li, Junda Wu, Jaehyeon Soon, Xiaofeng Zhang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

Comments (FinLLM 2023)@IJCAI 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.12180 2023-07-25 eess.IV cs.CV cs.LG 79%

Prototype-Driven and Multi-Expert Integrated Multi-Modal MR Brain Tumor Image Segmentation

Yafei Zhang, Zhiyuan Li, Huafeng Li, Dapeng Tao

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.09155 2023-07-19 cs.CV 79%

MLF-DET: Multi-Level Fusion for Cross-Modal 3D Object Detection

Zewei Lin, Yanqing Shen, Sanping Zhou, Shitao Chen, Nanning Zheng

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.08471 2023-07-18 cs.RO cs.AI 79%

Clarifying the Half Full or Half Empty Question: Multimodal Container Classification

Josua Spisak, Matthias Kerzel, Stefan Wermter

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments Preprint for ICANN 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.04907 2023-07-12 cs.CL cs.LG 79%

SimpleMTOD: A Simple Language Model for Multimodal Task-Oriented Dialogue with Symbolic Scene Representation

Bhathiya Hemanthage, Christian Dondrup, Phil Bartie, Oliver Lemon

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏