arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6918 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6918 篇

2312.05777 2023-12-18 cs.CV 79%

Negative Pre-aware for Noisy Cross-modal Matching

Xu Zhang, Hao Li, Mang Ye

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

Comments 9 pages, 5 figures, conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.08636 2023-12-15 cs.CV 79%

MmAP : Multi-modal Alignment Prompt for Cross-domain Multi-task Learning

Yi Xin, Junlong Du, Qiang Wang, Ke Yan, Shouhong Ding

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted by AAAI2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.08212 2023-12-14 cs.CV 79%

LAMM: Label Alignment for Multi-Modal Prompt Learning

Jingsheng Gao, Jiacheng Ruan, Suncheng Xiang, Zefang Yu, Ke Ji, Mingye Xie, Ting Liu, Yuzhuo Fu

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted at AAAI 2024 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.05434 2023-12-12 cs.CL 79%

Beneath the Surface: Unveiling Harmful Memes with Multimodal Reasoning Distilled from Large Language Models

Hongzhan Lin, Ziyang Luo, Jing Ma, Long Chen

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

Comments The first work to alleviate the issue of superficial understanding for harmful meme detection by explicitly utilizing commonsense knowledge, from a fresh perspective on harnessing advanced Large Language Models

Journal ref The 2023 Conference on Empirical Methods in Natural Language Processing

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.11312 2023-12-07 cs.CV 79%

Optimizing rgb-d semantic segmentation through multi-modal interaction and pooling attention

Shuai Zhang, Minghong Xie

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.02227 2023-12-06 cs.LG cs.CL 79%

Improving Multimodal Sentiment Analysis: Supervised Angular Margin-based Contrastive Learning for Enhanced Fusion Representation

Cong-Duy Nguyen, Thong Nguyen, Duc Anh Vu, Luu Anh Tuan

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.00360 2023-12-05 cs.CV 79%

Efficient Multimodal Semantic Segmentation via Dual-Prompt Learning

Shaohua Dong, Yunhe Feng, Qing Yang, Yan Huang, Dongfang Liu, Heng Fan

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments 11 pages, 4 figures, 9 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.00339 2023-12-05 cs.MM 79%

GraphMFT: A Graph Network based Multimodal Fusion Technique for Emotion Recognition in Conversation

Jiang Li, Xiaoping Wang, Guoqing Lv, Zhigang Zeng

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.MM

Comments Accepted by Neurocomputing

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.18245 2023-12-01 eess.IV cs.CV 79%

Automatic Detection of Alzheimer's Disease with Multi-Modal Fusion of Clinical MRI Scans

Long Chen, Liben Chen, Binfeng Xu, Wenxin Zhang, Narges Razavian

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.17597 2023-12-01 cs.CV 79%

Continual Self-supervised Learning: Towards Universal Multi-modal Medical Data Representation Learning

Yiwen Ye, Yutong Xie, Jianpeng Zhang, Ziyang Chen, Qi Wu, Yong Xia

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.01728 2023-12-01 cs.CV 79%

Generative-based Fusion Mechanism for Multi-Modal Tracking

Zhangyong Tang, Tianyang Xu, Xuefeng Zhu, Xiao-Jun Wu, Josef Kittler

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.01121 2023-11-23 cs.RO cs.CV 79%

Artifacts Mapping: Multi-Modal Semantic Mapping for Object Detection and 3D Localization

Federico Rollo, Gennaro Raiola, Andrea Zunino, Nikolaos Tsagarakis, Arash Ajoudani

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted to the 11th European Conference on Mobile Robots (ECMR) 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.11237 2023-11-21 cs.AI 79%

Implementation of AI Deep Learning Algorithm For Multi-Modal Sentiment Analysis

Jiazhen Wang

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.08032 2023-11-15 eess.IV cs.CV 79%

ELF: An End-to-end Local and Global Multimodal Fusion Framework for Glaucoma Grading

Wenyun Li, Chi-Man Pun

专题命中 多模态训练与对齐 :multimodal(title);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.07033 2023-11-14 eess.IV cs.CV 79%

TTMFN: Two-stream Transformer-based Multimodal Fusion Network for Survival Prediction

Ruiquan Ge, Xiangyang Hu, Rungen Huang, Gangyong Jia, Yaqi Wang, Renshu Gu, Changmiao Wang, Elazab Ahmed, Linyan Wang, Juan Ye, Ye Li

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.04058 2023-11-08 cs.CV 79%

mmFUSION: Multimodal Fusion for 3D Objects Detection

Javed Ahmad, Alessio Del Bue

专题命中 多模态训练与对齐 :multimodal(title);multi-modal(abstract);分类 cs.CV

Comments 10 pages,

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.03713 2023-11-08 quant-ph cs.CV cs.LG 79%

Multimodal deep representation learning for quantum cross-platform verification

Yang Qian, Yuxuan Du, Zhenliang He, Min-hsiu Hsieh, Dacheng Tao

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.18812 2023-10-31 cs.CV 79%

UniCat: Crafting a Stronger Fusion Baseline for Multimodal Re-Identification

Jennifer Crawford, Haoli Yin, Luke McDermott, Daniel Cummings

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Accepted NeurIPS 2023 UniReps, 9 pages, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.16936 2023-10-31 cs.CV cs.LG 79%

Diagnosing Alzheimer's Disease using Early-Late Multimodal Data Fusion with Jacobian Maps

Yasmine Mustafa, Tie Luo

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments To be published in Proceedings of 2023 IEEE Healthcom, December 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.16856 2023-10-27 cs.CV 79%

GraFT: Gradual Fusion Transformer for Multimodal Re-Identification

Haoli Yin, Jiayao Li, Eva Schiller, Luke McDermott, Daniel Cummings

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments 3 Borderline Reviews at WACV, 8 pages, 5 figures, 8 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.14496 2023-10-24 cs.MM 79%

Redundancy-Adaptive Multimodal Learning for Imperfect Data

Mengxi Chen, Jiangchao Yao, Linyu Xing, Yu Wang, Ya Zhang, Yanfeng Wang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.07871 2023-10-23 cs.AI 79%

Hierarchical Pretraining on Multimodal Electronic Health Records

Xiaochen Wang, Junyu Luo, Jiaqi Wang, Ziyi Yin, Suhan Cui, Yuan Zhong, Yaqing Wang, Fenglong Ma

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments Accepted by EMNLP 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.12802 2023-10-23 cs.LG cs.AI q-bio.BM 79%

Otter-Knowledge: benchmarks of multimodal knowledge graph representation learning from different sources for drug discovery

Hoang Thanh Lam, Marco Luca Sbodio, Marcos Martínez Galindo, Mykhaylo Zayats, Raúl Fernández-Díaz, Víctor Valls, Gabriele Picco, Cesar Berrospi Ramis, Vanessa López

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.11910 2023-10-19 eess.IV cs.CV cs.LG 79%

Multi-modal Medical Neurological Image Fusion using Wavelet Pooled Edge Preserving Autoencoder

Manisha Das, Deep Gupta, Petia Radeva, Ashwini M Bakde

专题命中 多模态训练与对齐 :multi-modal(title);multimodal(abstract);分类 cs.CV

Comments 8 pages, 5 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.11896 2023-10-19 eess.IV cs.CV cs.LG 79%

A New Multimodal Medical Image Fusion based on Laplacian Autoencoder with Channel Attention

Payal Wankhede, Manisha Das, Deep Gupta, Petia Radeva, Ashwini M Bakde

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments 10 pages, 6 figures, % tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.10649 2023-10-17 cs.CV 79%

Cross-modal and Cross-domain Knowledge Transfer for Label-free 3D Segmentation

Jingyu Zhang, Huitong Yang, Dai-Jie Wu, Jacky Keung, Xuesong Li, Xinge Zhu, Yuexin Ma

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

Comments 12 pages,4 figures,accepted

Journal ref Chinese Conference on Pattern Recognition and Computer Vision (PRCV) 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.06365 2023-10-11 cs.CL 79%

Multi-Modal Knowledge Graph Transformer Framework for Multi-Modal Entity Alignment

Qian Li, Cheng Ji, Shu Guo, Zhaoji Liang, Lihong Wang, Jianxin Li

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.03485 2023-10-10 eess.IV cs.CV cs.LG 79%

BTDNet: a Multi-Modal Approach for Brain Tumor Radiogenomic Classification

Dimitrios Kollias, Karanjot Vendal, Priyanka Gadhavi, Solomon Russom

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.08747 2023-10-06 cs.CV 79%

Unified Brain MR-Ultrasound Synthesis using Multi-Modal Hierarchical Representations

Reuben Dorent, Nazim Haouchine, Fryderyk Kögl, Samuel Joutard, Parikshit Juvekar, Erickson Torio, Alexandra Golby, Sebastien Ourselin, Sarah Frisken, Tom Vercauteren, Tina Kapur, William M. Wells

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted at MICCAI 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.02960 2023-10-05 cs.CV 79%

CoDA: Collaborative Novel Box Discovery and Cross-modal Alignment for Open-vocabulary 3D Object Detection

Yang Cao, Yihan Zeng, Hang Xu, Dan Xu

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

Comments Accepted by NeurIPS 2023. Project Page: https://yangcaoai.github.io/publications/CoDA.html

详情

展开后加载摘要…

URL PDF HTML 收藏