arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6918 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6918 篇

2404.13619 2024-04-23 cs.MM 79%

Towards Unified Representation of Multi-Modal Pre-training for 3D Understanding via Differentiable Rendering

Ben Fei, Yixuan Li, Weidong Yang, Lipeng Ma, Ying He

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.12588 2024-04-22 cs.CV cs.LG 79%

Cross-Modal Adapter: Parameter-Efficient Transfer Learning Approach for Vision-Language Models

Juncheng Yang, Zuchao Li, Shuai Xie, Weiping Zhu, Wei Yu, Shijun Li

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

Comments This paper is accepted to ICME 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.16666 2024-04-22 cs.LG cs.AI physics.chem-ph q-bio.BM 79%

MultiModal-Learning for Predicting Molecular Properties: A Framework Based on Image and Graph Structures

Zhuoyuan Wang, Jiacong Mi, Shan Lu, Jieyue He

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.05927 2024-04-22 cs.LG cs.AI eess.SP 79%

Frequency-Aware Masked Autoencoders for Multimodal Pretraining on Biosignals

Ran Liu, Ellen L. Zippi, Hadi Pouransari, Chris Sandino, Jingping Nie, Hanlin Goh, Erdrin Azemi, Ali Moin

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments Extended version of ICLR 2024 Learning from Time Series for Health workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.11764 2024-04-19 cs.CV 79%

Multimodal 3D Object Detection on Unseen Domains

Deepti Hegde, Suhas Lohit, Kuan-Chuan Peng, Michael J. Jones, Vishal M. Patel

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments technical report

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.10381 2024-04-18 cs.IR cs.AI 79%

UMAIR-FPS: User-aware Multi-modal Animation Illustration Recommendation Fusion with Painting Style

Yan Kang, Hao Lin, Mingjian Yang, Shin-Jye Lee

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.AI

Comments Accepted by DASFAA 2024 Research track

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.10146 2024-04-17 cs.CV 79%

Cross-Modal Self-Training: Aligning Images and Pointclouds to Learn Classification without Labels

Amaya Dharmasiri, Muzammal Naseer, Salman Khan, Fahad Shahbaz Khan

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

Comments To be published in Workshop for Learning 3D with Multi-View Supervision (3DMV) at CVPR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.09499 2024-04-16 cs.CV cs.GR 79%

Learning Human Motion from Monocular Videos via Cross-Modal Manifold Alignment

Shuaiying Hou, Hongyu Tao, Junheng Fang, Changqing Zou, Hujun Bao, Weiwei Xu

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.08923 2024-04-16 cs.CV 79%

Trustworthy Multimodal Fusion for Sentiment Analysis in Ordinal Sentiment Space

Zhuyang Xie, Yan Yang, Jie Wang, Xiaorong Liu, Xiaofan Li

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments 14 pages, 9 figures, Accepted by IEEE Transactions on Circuits and Systems for Video Technology

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.06791 2024-04-16 cs.CV 79%

PV-SSD: A Multi-Modal Point Cloud Feature Fusion Method for Projection Features and Variable Receptive Field Voxel Features

Yongxin Shao, Aihong Tan, Zhetao Sun, Enhui Zheng, Tianhong Yan, Peng Liao

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.06107 2024-04-10 cs.CL 79%

Exploring the Necessity of Visual Modality in Multimodal Machine Translation using Authentic Datasets

Zi Long, Zhenhao Tang, Xianghua Fu, Jian Chen, Shilong Hou, Jinze Lyu

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

Comments bucc 2024 accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.16119 2024-04-09 cs.AI 79%

Triple Disentangled Representation Learning for Multimodal Affective Analysis

Ying Zhou, Xuefeng Liang, Han Chen, Yin Zhao, Xin Chen, Lida Yu

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments 14 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.04026 2024-04-08 cs.RO cs.CV 79%

MM-Gaussian: 3D Gaussian-based Multi-modal Fusion for Localization and Reconstruction in Unbounded Scenes

Chenyang Wu, Yifan Duan, Xinran Zhang, Yu Sheng, Jianmin Ji, Yanyong Zhang

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments 7 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.00144 2024-04-02 eess.IV cs.CV 79%

An Interpretable Cross-Attentive Multi-modal MRI Fusion Framework for Schizophrenia Diagnosis

Ziyu Zhou, Anton Orlichenko, Gang Qu, Zening Fu, Vince D Calhoun, Zhengming Ding, Yu-Ping Wang

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.16002 2024-03-29 cs.CV 79%

SDSTrack: Self-Distillation Symmetric Adapter Learning for Multi-Modal Visual Object Tracking

Xiaojun Hou, Jiazheng Xing, Yijie Qian, Yaowei Guo, Shuo Xin, Junhao Chen, Kai Tang, Mengmeng Wang, Zhengkai Jiang, Liang Liu, Yong Liu

专题命中 多模态训练与对齐 :multi-modal(title);multimodal(abstract);分类 cs.CV

Comments Accepted by CVPR2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.16257 2024-03-26 cs.CV 79%

Unlearning Backdoor Threats: Enhancing Backdoor Defense in Multimodal Contrastive Learning via Local Token Unlearning

Siyuan Liang, Kuanrong Liu, Jiajun Gong, Jiawei Liang, Yuan Xun, Ee-Chien Chang, Xiaochun Cao

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments 6 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.01099 2024-03-26 eess.IV cs.CV cs.LG 79%

HyMNet: a Multimodal Deep Learning System for Hypertension Classification using Fundus Photographs and Cardiometabolic Risk Factors

Mohammed Baharoon, Hessa Almatar, Reema Alduhayan, Tariq Aldebasi, Badr Alahmadi, Yahya Bokhari, Mohammed Alawad, Ahmed Almazroa, Abdulrhman Aljouie

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.15241 2024-03-25 cs.CV 79%

IS-Fusion: Instance-Scene Collaborative Fusion for Multimodal 3D Object Detection

Junbo Yin, Jianbing Shen, Runnan Chen, Wei Li, Ruigang Yang, Pascal Frossard, Wenguan Wang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Accepted to CVPR 2024; Code: https://github.com/yinjunbo/IS-Fusion

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.08709 2024-03-25 cs.CV 79%

You Only Need Two Detectors to Achieve Multi-Modal 3D Multi-Object Tracking

Xiyang Wang, Chunyun Fu, Jiawei He, Mingguang Huang, Ting Meng, Siyu Zhang, Hangning Zhou, Ziyao Xu, Chi Zhang

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments 11 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.13830 2024-03-22 q-bio.BM cs.CL cs.LG 79%

Bridging Text and Molecule: A Survey on Multimodal Frameworks for Molecule

Yi Xiao, Xiangxin Zhou, Qiang Liu, Liang Wang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.12725 2024-03-21 cs.CL 79%

Generative Multimodal Entity Linking

Senbao Shi, Zhenran Xu, Baotian Hu, Min Zhang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

Comments Accepted by LREC-COLING 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.10036 2024-03-18 cs.CV 79%

SparseFusion: Efficient Sparse Multi-Modal Fusion Framework for Long-Range 3D Perception

Yiheng Li, Hongyang Li, Zehao Huang, Hong Chang, Naiyan Wang

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.08182 2024-03-14 cs.CV 79%

SeCG: Semantic-Enhanced 3D Visual Grounding via Cross-modal Graph Attention

Feng Xiao, Hongbin Xu, Qiuxia Wu, Wenxiong Kang

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.04829 2024-03-14 cs.CV 79%

MixReorg: Cross-Modal Mixed Patch Reorganization is a Good Mask Learner for Open-World Semantic Segmentation

Kaixin Cai, Pengzhen Ren, Yi Zhu, Hang Xu, Jianzhuang Liu, Changlin Li, Guangrun Wang, Xiaodan Liang

专题命中 多模态训练与对齐 :cross-modal(title);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.07301 2024-03-13 cs.CV 79%

Let Storytelling Tell Vivid Stories: An Expressive and Fluent Multimodal Storyteller

Chuanqi Zang, Jiji Tang, Rongsheng Zhang, Zeng Zhao, Tangjie Lv, Mingtao Pei, Wei Liang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.10601 2024-03-13 cs.CV eess.SP 79%

Multimodal Indoor Localization Using Crowdsourced Radio Maps

Zhaoguang Yi, Xiangyu Wen, Qiyue Xia, Peize Li, Francisco Zampella, Firas Alsehly, Chris Xiaoxuan Lu

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments 7 pages, 4 figures; ICRA'24 https://youtu.be/NTTKwJBFN5w

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.05552 2024-03-12 cs.CY cs.AI cs.LG 79%

Multi-source and multimodal data fusion for predicting academic performance in blended learning university courses

W. Chango, R. Cerezo, C. Romero

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Journal ref Computers & Electrical Engineering, 89, 106908 (2021)

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.06339 2024-03-12 cs.CV 79%

FOAA: Flattened Outer Arithmetic Attention For Multimodal Tumor Classification

Omnia Alwazzan, Ioannis Patras, Gregory Slabaugh

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments This paper has been accepted for ISBI-2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.01802 2024-03-12 cs.CV 79%

TNF: Tri-branch Neural Fusion for Multimodal Medical Data Classification

Tong Zheng, Shusaku Sone, Yoshitaka Ushiku, Yuki Oba, Jiaxin Ma

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.03217 2024-03-06 cs.CV 79%

Self-supervised 3D Patient Modeling with Multi-modal Attentive Fusion

Meng Zheng, Benjamin Planche, Xuan Gong, Fan Yang, Terrence Chen, Ziyan Wu

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments MICCAI 2022

详情

展开后加载摘要…

URL PDF HTML 收藏