arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6903 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6903 篇

2406.09696 2024-06-17 eess.IV cs.CV 83%

MoME: Mixture of Multimodal Experts for Cancer Survival Prediction

Conghao Xiong, Hao Chen, Hao Zheng, Dong Wei, Yefeng Zheng, Joseph J. Y. Sung, Irwin King

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments 8 + 1/2 pages, early accepted to MICCAI2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.05130 2024-06-10 cs.CL 83%

An Empirical Study on Parameter-Efficient Fine-Tuning for MultiModal Large Language Models

Xiongtao Zhou, Jie He, Yuhua Ke, Guangyao Zhu, Víctor Gutiérrez-Basulto, Jeff Z. Pan

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CL

Comments ACL finding 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.02263 2024-06-05 cs.CV 83%

M3DM-NR: RGB-3D Noisy-Resistant Industrial Anomaly Detection via Multimodal Denoising

Chengjie Wang, Haokun Zhu, Jinlong Peng, Yue Wang, Ran Yi, Yunsheng Wu, Lizhuang Ma, Jiangning Zhang

专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.01210 2024-06-05 cs.CV 83%

GeminiFusion: Efficient Pixel-wise Multimodal Fusion for Vision Transformer

Ding Jia, Jianyuan Guo, Kai Han, Han Wu, Chao Zhang, Chang Xu, Xinghao Chen

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted by ICML 2024, code and models are available at https://github.com/JiaDingCN/GeminiFusion

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.20985 2024-06-03 cs.CV 83%

DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models

Linli Yao, Lei Li, Shuhuai Ren, Lean Wang, Yuanxin Liu, Xu Sun, Lu Hou

专题命中 多模态训练与对齐 :multimodal(title);MLLM(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.16556 2024-04-25 cs.LG cs.AI 83%

LANISTR: Multimodal Learning from Structured and Unstructured Data

Sayna Ebrahimi, Sercan O. Arik, Yihe Dong, Tomas Pfister

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.04001 2024-04-22 cs.CV cs.LG 83%

MMSFormer: Multimodal Transformer for Material and Semantic Segmentation

Md Kaykobad Reza, Ashley Prater-Bennette, M. Salman Asif

专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments Accepted by IEEE Open Journal of Signal Processing. 15 pages, 3 figures, 9 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.17187 2024-04-18 eess.IV cs.CV 83%

PE-MVCNet: Multi-view and Cross-modal Fusion Network for Pulmonary Embolism Prediction

Zhaoxin Guo, Zhipeng Wang, Ruiquan Ge, Jianxun Yu, Feiwei Qin, Yuan Tian, Yuqing Peng, Yonghong Li, Changmiao Wang

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.15432 2024-04-16 cs.CL 83%

CFN-ESA: A Cross-Modal Fusion Network with Emotion-Shift Awareness for Dialogue Emotion Recognition

Jiang Li, Xiaoping Wang, Yingjian Liu, Zhigang Zeng

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CL

Comments Accepted by IEEE Transactions on Affective Computing (TAFFC)

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.10707 2024-04-02 cs.LG cs.CV 83%

Multimodal Representation Learning by Alternating Unimodal Adaptation

Xiaohui Zhang, Jaehong Yoon, Mohit Bansal, Huaxiu Yao

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted by CVPR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.16558 2024-04-01 cs.CV 83%

Elysium: Exploring Object-level Perception in Videos via MLLM

Han Wang, Yanjie Wang, Yongjie Ye, Yuxiang Nie, Can Huang

专题命中 多模态训练与对齐 :MLLM(title,abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.19203 2024-03-30 eess.IV cs.CV 83%

Single-Shared Network with Prior-Inspired Loss for Parameter-Efficient Multi-Modal Imaging Skin Lesion Classification

Peng Tang, Tobias Lasser

专题命中 多模态训练与对齐 :multi-modal(title,abstract);multimodal(abstract);分类 cs.CV

Comments This paper have submitted to Journal for review

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.14565 2024-03-21 cs.CV cs.AI cs.CE cs.CL cs.MM 83%

Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning

Fuxiao Liu, Kevin Lin, Linjie Li, Jianfeng Wang, Yaser Yacoob, Lijuan Wang

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments 40 pages, 32 figures, ICLR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.09672 2024-03-18 cs.CV cs.LG 83%

COMPRER: A Multimodal Multi-Objective Pretraining Framework for Enhanced Medical Image Representation

Guy Lutsker, Hagai Rossman, Nastya Godiva, Eran Segal

专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.07372 2024-03-13 cs.CV 83%

Eliminating Cross-modal Conflicts in BEV Space for LiDAR-Camera 3D Object Detection

Jiahui Fu, Chen Gao, Zitian Wang, Lirong Yang, Xiaofei Wang, Beipeng Mu, Si Liu

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments Accepted by ICRA 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.17911 2024-03-13 cs.CV 83%

OPERA: Alleviating Hallucination in Multi-Modal Large Language Models via Over-Trust Penalty and Retrospection-Allocation

Qidong Huang, Xiaoyi Dong, Pan Zhang, Bin Wang, Conghui He, Jiaqi Wang, Dahua Lin, Weiming Zhang, Nenghai Yu

专题命中 多模态训练与对齐 :multi-modal(title,abstract);MLLM(abstract);分类 cs.CV

Comments CVPR 2024, code is available at https://github.com/shikiw/OPERA

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.03707 2024-03-07 cs.CV 83%

Multi-Grained Cross-modal Alignment for Learning Open-vocabulary Semantic Segmentation from Text Supervision

Yajie Liu, Pu Ge, Qingjie Liu, Di Huang

专题命中 多模态训练与对齐 :cross-modal(title,abstract);image-text(abstract);分类 cs.CV

Comments 17 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.19298 2024-03-06 cs.CV 83%

Suppress and Rebalance: Towards Generalized Multi-Modal Face Anti-Spoofing

Xun Lin, Shuai Wang, Rizhao Cai, Yizhong Liu, Ying Fu, Zitong Yu, Wenzhong Tang, Alex Kot

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments Accepeted by CVPR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.15858 2024-02-27 cs.CV cs.DC 83%

FedMM: Federated Multi-Modal Learning with Modality Heterogeneity in Computational Pathology

Yuanzhe Peng, Jieming Bian, Jie Xu

专题命中 多模态训练与对齐 :multi-modal(title,abstract);multimodal(abstract);分类 cs.CV

Journal ref 2024 International Conference on Acoustics, Speech and Signal Processing (ICASSP 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.09503 2024-01-26 cs.CV 83%

JM3D & JM3D-LLM: Elevating 3D Understanding with Joint Multi-modal Cues

Jiayi Ji, Haowei Wang, Changli Wu, Yiwei Ma, Xiaoshuai Sun, Rongrong Ji

专题命中 多模态训练与对齐 :multi-modal(title,abstract);multimodal(abstract);分类 cs.CV

Comments 16 pages, 4 figures, 10 tables, 3D understanding

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.02982 2024-01-26 cs.CV 83%

Beyond First Impressions: Integrating Joint Multi-modal Cues for Comprehensive 3D Representation

Haowei Wang, Jiji Tang, Jiayi Ji, Xiaoshuai Sun, Rongsheng Zhang, Yiwei Ma, Minda Zhao, Lincheng Li, zeng zhao, Tangjie Lv, Rongrong Ji

专题命中 多模态训练与对齐 :multi-modal(title,abstract);multimodal(abstract);分类 cs.CV

Comments ACM MM 2023, 3D Understanding, JM3D

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.09638 2024-01-19 eess.IV cs.CV cs.LG 83%

Automatic 3D Multi-modal Ultrasound Segmentation of Human Placenta using Fusion Strategies and Deep Learning

Sonit Singh, Gordon Stevenson, Brendan Mein, Alec Welsh, Arcot Sowmya

专题命中 多模态训练与对齐 :multi-modal(title,abstract);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.08123 2024-01-17 cs.CV 83%

The Devil is in the Details: Boosting Guided Depth Super-Resolution via Rethinking Cross-Modal Alignment and Aggregation

Xinni Jiang, Zengsheng Kuang, Chunle Guo, Ruixun Zhang, Lei Cai, Xiao Fan, Chongyi Li

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.16818 2024-01-09 cs.CV 83%

SimDistill: Simulated Multi-modal Distillation for BEV 3D Object Detection

Haimei Zhao, Qiming Zhang, Shanshan Zhao, Zhe Chen, Jing Zhang, Dacheng Tao

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted by AAAI 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.16648 2023-12-29 cs.RO cs.CV 83%

LIP-Loc: LiDAR Image Pretraining for Cross-Modal Localization

Sai Shubodh Puligilla, Mohammad Omama, Husain Zaidi, Udit Singh Parihar, Madhava Krishna

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments To be presented at WACV-W 2024. Project page: https://shubodhs.ai/liploc

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.15645 2023-12-27 cs.CL 83%

Conditional Variational Autoencoder for Sign Language Translation with Cross-Modal Alignment

Rui Zhao, Liang Zhang, Biao Fu, Cong Hu, Jinsong Su, Yidong Chen

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CL

Comments Accepted as conference paper by AAAI24. The code and models are available at https://github.com/rzhao-zhsq/CV-SLT

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.13929 2023-12-27 cs.CV 83%

XKD: Cross-modal Knowledge Distillation with Domain Alignment for Video Representation Learning

Pritam Sarkar, Ali Etemad

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV

Comments AAAI 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.14471 2023-12-25 cs.CV 83%

Prototype-based Cross-Modal Object Tracking

Lei Liu, Chenglong Li, Futian Wang, Longfeng Shen, Jin Tang

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments In Peer Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.14446 2023-12-25 cs.CV 83%

Cross-Modal Object Tracking via Modality-Aware Fusion Network and A Large-Scale Dataset

Lei Liu, Mengya Zhang, Cheng Li, Chenglong Li, Jin Tang

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments In Peer Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.05804 2023-12-15 cs.AI cs.CL cs.CV cs.MM 83%

Learning Language-guided Adaptive Hyper-modality Representation for Multimodal Sentiment Analysis

Haoyu Zhang, Yu Wang, Guanghao Yin, Kejun Liu, Yuanyuan Liu, Tianshu Yu

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Published in EMNLP 2023

Journal ref Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing

详情

展开后加载摘要…

URL PDF HTML 收藏