arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6918 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6918 篇

2403.02991 2024-03-06 cs.CV 79%

MADTP: Multimodal Alignment-Guided Dynamic Token Pruning for Accelerating Vision-Language Transformer

Jianjian Cao, Peng Ye, Shengze Li, Chong Yu, Yansong Tang, Jiwen Lu, Tao Chen

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments 19 pages, 9 figures, Published in CVPR2024

Journal ref In Proc. IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.01203 2024-03-05 cs.LG cs.CL cs.DB 79%

Pseudo-Label Calibration Semi-supervised Multi-Modal Entity Alignment

Luyao Wang, Pengnian Qi, Xigang Bao, Chunlai Zhou, Biao Qin

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CL

Comments accepted by AAAI2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.17936 2024-02-29 cs.CL 79%

Acquiring Linguistic Knowledge from Multimodal Input

Theodor Amariucai, Alex Warstadt

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

Comments in Proceedings of the BabyLM Challenge at the 27th Conference on Computational Natural Language Learning

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.09831 2024-02-29 eess.IV cs.CV 79%

Cross-modality Attention-based Multimodal Fusion for Non-small Cell Lung Cancer (NSCLC) Patient Survival Prediction

Ruining Deng, Nazim Shaikh, Gareth Shannon, Yao Nie

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.17483 2024-02-28 cs.CV 79%

AlignMiF: Geometry-Aligned Multimodal Implicit Field for LiDAR-Camera Joint Synthesis

Tao Tang, Guangrun Wang, Yixing Lao, Peng Chen, Jie Liu, Liang Lin, Kaicheng Yu, Xiaodan Liang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments CVPR2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.11826 2024-02-20 cs.CV 79%

Unveiling the Depths: A Multi-Modal Fusion Framework for Challenging Scenarios

Jialei Xu, Xianming Liu, Junjun Jiang, Kui Jiang, Rui Li, Kai Cheng, Xiangyang Ji

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.11307 2024-02-20 cs.CV 79%

ICHPro: Intracerebral Hemorrhage Prognosis Classification Via Joint-attention Fusion-based 3d Cross-modal Network

Xinlei Yu, Xinyang Li, Ruiquan Ge, Shibin Wu, Ahmed Elazab, Jichao Zhu, Lingyan Zhang, Gangyong Jia, Taosheng Xu, Xiang Wan, Changmiao Wang

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

Comments 6 pages,4 figures, 4 tables, accepted by ISBI

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.07358 2024-02-19 cs.CL 79%

Towards Versatile and Efficient Visual Knowledge Integration into Pre-trained Language Models with Cross-Modal Adapters

Xinyun Zhang, Haochen Tan, Han Wu, Bei Yu

专题命中 多模态训练与对齐 :cross-modal(title);multi-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.09738 2024-02-16 cs.CL 79%

Align before Attend: Aligning Visual and Textual Features for Multimodal Hateful Content Detection

Eftekhar Hossain, Omar Sharif, Mohammed Moshiul Hoque, Sarah M. Preum

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

Comments Accepted to EACL-SRW, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.02384 2024-02-16 cs.CV 79%

ChartAssisstant: A Universal Chart Multimodal Language Model via Chart-to-Table Pre-training and Multitask Instruction Tuning

Fanqing Meng, Wenqi Shao, Quanfeng Lu, Peng Gao, Kaipeng Zhang, Yu Qiao, Ping Luo

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Updated and corrected experimental results, removal of inappropriate experiments, and a more comprehensive experimental setup

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.05151 2024-02-09 cs.LG cs.AI 79%

CrashFormer: A Multimodal Architecture to Predict the Risk of Crash

Amin Karimi Monsefi, Pouya Shiri, Ahmad Mohammadshirazi, Nastaran Karimi Monsefi, Ron Davies, Sobhan Moosavi, Rajiv Ramnath

专题命中 多模态训练与对齐 :multimodal(title);multi-modal(abstract);分类 cs.AI

Comments The paper is accepted In 1st ACM SIGSPATIAL International Workshop on Advances in Urban-AI (UrbanAI 23), November 13, 2023, Hamburg, Germany

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.02055 2024-02-06 cs.LG cs.AI 79%

Variance Alignment Score: A Simple But Tough-to-Beat Data Selection Method for Multimodal Contrastive Learning

Yiping Wang, Yifang Chen, Wendan Yan, Kevin Jamieson, Simon Shaolei Du

专题命中 多模态训练与对齐 :multimodal(title);multi-modal(abstract);分类 cs.AI

Comments 17 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.13508 2024-02-06 cs.LG cs.AI cs.DC 79%

Multimodal Federated Learning with Missing Modality via Prototype Mask and Contrast

Guangyin Bao, Qi Zhang, Duoqian Miao, Zixuan Gong, Liang Hu, Ke Liu, Yang Liu, Chongyang Shi

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments 23 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.01311 2024-02-05 cs.CV eess.IV 79%

Deep Multimodal Fusion of Data with Heterogeneous Dimensionality via Projective Networks

José Morano, Guilherme Aresta, Christoph Grechenig, Ursula Schmidt-Erfurth, Hrvoje Bogunović

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Accepted for publication in the IEEE Journal of Biomedical and Health Informatics (JBHI)

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.01886 2024-02-01 cs.CV 79%

Bridging the Gap between Multi-focus and Multi-modal: A Focused Integration Framework for Multi-modal Image Fusion

Xilai Li, Xiaosong Li, Tao Ye, Xiaoqi Cheng, Wuyang Liu, Haishu Tan

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted to IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.09479 2024-01-24 cs.CR cs.AI cs.LG 79%

Uncertainty-Aware Hardware Trojan Detection Using Multimodal Deep Learning

Rahul Vishwakarma, Amin Rezaei

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments 2024 Design, Automation and Test in Europe Conference | The European Event for Electronic System Design & Test (accepted)

Journal ref 2024 Design, Automation and Test in Europe Conference | The European Event for Electronic System Design & Test

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.03179 2024-01-24 cs.CV 79%

Multimodal Informative ViT: Information Aggregation and Distribution for Hyperspectral and LiDAR Classification

Jiaqing Zhang, Jie Lei, Weiying Xie, Geng Yang, Daixun Li, Yunsong Li

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.11740 2024-01-23 cs.CV cs.LG 79%

Multi-level Cross-modal Alignment for Image Clustering

Liping Qiu, Qin Zhang, Xiaojun Chen, Shaotian Cai

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.11252 2024-01-23 cs.LG cs.AI 79%

Automated Fusion of Multimodal Electronic Health Records for Better Medical Predictions

Suhan Cui, Jiaqi Wang, Yuan Zhong, Han Liu, Ting Wang, Fenglong Ma

专题命中 多模态训练与对齐 :multimodal(title);multi-modal(abstract);分类 cs.AI

Comments Accepted by SDM 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.02332 2024-01-23 cs.LG cs.CV 79%

Multimodal Machine Learning in Image-Based and Clinical Biomedicine: Survey and Prospects

Elisa Warner, Joonsang Lee, William Hsu, Tanveer Syeda-Mahmood, Charles Kahn, Olivier Gevaert, Arvind Rao

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.10041 2024-01-19 cs.CV 79%

CMFN: Cross-Modal Fusion Network for Irregular Scene Text Recognition

Jinzhi Zheng, Ruyi Ji, Libo Zhang, Yanjun Wu, Chen Zhao

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

Comments Accepted to ICONIP 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.07854 2024-01-17 cs.CV 79%

$M^{2}$Fusion: Bayesian-based Multimodal Multi-level Fusion on Colorectal Cancer Microsatellite Instability Prediction

Quan Liu, Jiawen Yao, Lisha Yao, Xin Chen, Jingren Zhou, Le Lu, Ling Zhang, Zaiyi Liu, Yuankai Huo

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.07571 2024-01-17 cs.CV 79%

A Bi-Pyramid Multimodal Fusion Method for the Diagnosis of Bipolar Disorders

Guoxin Wang, Sheng Shi, Shan An, Fengmei Fan, Wenshu Ge, Qi Wang, Feng Yu, Zhiren Wang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Accepted by IEEE ICASSP 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.03082 2024-01-09 cs.AI 79%

UMIE: Unified Multimodal Information Extraction with Instruction Tuning

Lin Sun, Kai Zhang, Qingyuan Li, Renze Lou

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.01364 2024-01-04 q-bio.NC cs.AI cs.LG cs.NE 79%

Multi-Modal Cognitive Maps based on Neural Networks trained on Successor Representations

Paul Stoewer, Achim Schilling, Andreas Maier, Patrick Krauss

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.08965 2024-01-03 cs.CV 79%

PointDC:Unsupervised Semantic Segmentation of 3D Point Clouds via Cross-modal Distillation and Super-Voxel Clustering

Zisheng Chen, Hongbin Xu, Weitao Chen, Zhipeng Zhou, Haihong Xiao, Baigui Sun, Xuansong Xie, Wenxiong Kang

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

Comments Accepted by International Conference on Computer Vision (ICCV) 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.00268 2024-01-02 cs.CV 79%

COMMA: Co-Articulated Multi-Modal Learning

Lianyu Hu, Liqing Gao, Zekang Liu, Chi-Man Pun, Wei Feng

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted to AAAI2024. Code is available at https://github.com/hulianyuyy/COMMA

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.16797 2023-12-29 cs.CV 79%

Multi-Prompts Learning with Cross-Modal Alignment for Attribute-based Person Re-Identification

Yajing Zhai, Yawen Zeng, Zhiyong Huang, Zheng Qin, Xin Jin, Da Cao

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

Comments AAAI 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.16279 2023-12-29 cs.CV 79%

Cloud-Device Collaborative Learning for Multimodal Large Language Models

Guanqun Wang, Jiaming Liu, Chenxuan Li, Junpeng Ma, Yuan Zhang, Xinyu Wei, Kevin Zhang, Maurice Chong, Ray Zhang, Yijiang Liu, Shanghang Zhang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.14410 2023-12-25 cs.CV 79%

A Multi-Stage Adaptive Feature Fusion Neural Network for Multimodal Gait Recognition

Shinan Zou, Jianbo Xiong, Chao Fan, Shiqi Yu, Jin Tang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments This paper has been accepted by IJCB2023

Journal ref IJCB2023

详情

展开后加载摘要…

URL PDF HTML 收藏