arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6903 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6903 篇

2109.05178 2021-09-14 cs.CL 83%

College Student Retention Risk Analysis From Educational Database using Multi-Task Multi-Modal Neural Fusion

Mohammad Arif Ul Alam

专题命中 多模态训练与对齐 :multi-modal(title,abstract);multimodal(abstract);分类 cs.CL

Comments Submitted to 36th AAAI Conference on Artificial Intelligence (AAAI) 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.12899 2021-08-31 cs.CL cs.LG 83%

Fine-Grained Chemical Entity Typing with Multimodal Knowledge Representation

Chenkai Sun, Weijiang Li, Jinfeng Xiao, Nikolaus Nova Parulian, ChengXiang Zhai, Heng Ji

专题命中 多模态训练与对齐 :multimodal(title);multi-modal(abstract);cross-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.12449 2021-08-11 cs.CV 83%

FusionPainting: Multimodal Fusion with Adaptive Attention for 3D Object Detection

Shaoqing Xu, Dingfu Zhou, Jin Fang, Junbo Yin, Zhou Bin, Liangjun Zhang

专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments Accepted by ITSC 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.06652 2021-08-09 cs.CV 83%

Self-Supervised Multi-Modal Alignment for Whole Body Medical Imaging

Rhydian Windsor, Amir Jamaludin, Timor Kadir, Andrew Zisserman

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted as a full paper to MICCAI 2021. Code will be made publicly available before September 27th 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.00206 2021-07-02 cs.LG cs.CV 83%

Multi-modal Graph Learning for Disease Prediction

Shuai Zheng, Zhenfeng Zhu, Zhizhe Liu, Zhenyu Guo, Yang Liu, Yao Zhao

专题命中 多模态训练与对齐 :multi-modal(title,abstract);multimodal(abstract);分类 cs.CV

Comments 10 pages, 4 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.05438 2021-06-11 cs.CV 83%

Cross-Modal Discrete Representation Learning

Alexander H. Liu, SouYoung Jin, Cheng-I Jeff Lai, Andrew Rouditchenko, Aude Oliva, James Glass

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.01631 2021-03-10 cs.LG cs.MM 83%

Robust Latent Representations via Cross-Modal Translation and Alignment

Vandana Rajan, Alessio Brutti, Andrea Cavallaro

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.MM

Journal ref ICASSP 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.11211 2020-12-22 eess.IV cs.CV 83%

A Multi-View Dynamic Fusion Framework: How to Improve the Multimodal Brain Tumor Segmentation from Multi-Views?

Yi Ding, Wei Zheng, Guozheng Wu, Ji Geng, Mingsheng Cao, Zhiguang Qin

专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments 21 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.03908 2020-11-10 eess.IV cs.CV cs.LG 83%

Cross-Modal Self-Attention Distillation for Prostate Cancer Segmentation

Guokai Zhang, Xiaoang Shen, Ye Luo, Jihao Luo, Zeju Wang, Weigang Wang, Binghui Zhao, Jianwei Lu

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments 2020 IEEE International Conference on Bioinformatics and Biomedicine (BIBM)

详情

展开后加载摘要…

URL PDF HTML 收藏
1912.05396 2020-10-27 cs.CV cs.LG 83%

Multimodal Self-Supervised Learning for Medical Image Analysis

Aiham Taleb, Christoph Lippert, Tassilo Klein, Moin Nabi

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments NeurIPS 2019 Workshops

详情

展开后加载摘要…

URL PDF HTML 收藏
2009.01523 2020-09-04 cs.CV 83%

A Comparison of Pre-trained Vision-and-Language Models for Multimodal Representation Learning across Medical Images and Reports

Yikuan Li, Hanyin Wang, Yuan Luo

专题命中 多模态训练与对齐 :multimodal(title,abstract);image-text(abstract);分类 cs.CV

Comments 10 pages, 3 figures, submitted to BIBM2020

详情

展开后加载摘要…

URL PDF HTML 收藏
1907.00135 2019-09-17 cs.CV 83%

RFBNet: Deep Multimodal Networks with Residual Fusion Blocks for RGB-D Semantic Segmentation

Liuyuan Deng, Ming Yang, Tianyi Li, Yuesheng He, Chunxiang Wang

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1908.05649 2019-08-16 eess.IV cs.CV 83%

A Multimodal Vision Sensor for Autonomous Driving

Dongming Sun, Xiao Huang, Kailun Yang

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1904.13072 2019-05-01 cs.CV 83%

Cross-Modal Message Passing for Two-stream Fusion

Dong Wang, Yuan Yuan, Qi Wang

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments 2018 IEEE International Conference on Acoustics, Speech and Signal Processing

详情

展开后加载摘要…

URL PDF HTML 收藏
1903.06496 2019-03-18 cs.LG cs.CV cs.NE 83%

MFAS: Multimodal Fusion Architecture Search

Juan-Manuel Pérez-Rúa, Valentin Vielzeuf, Stéphane Pateux, Moez Baccouche, Frédéric Jurie

专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments CVPR 2019, Jun 2019, Long Beach, United States http://cvpr2019.thecvf.com/

详情

展开后加载摘要…

URL PDF HTML 收藏
1711.00049 2018-06-19 cs.CV cs.LG 83%

Medical Image Segmentation Based on Multi-Modal Convolutional Neural Network: Study on Image Fusion Schemes

Zhe Guo, Xiang Li, Heng Huang, Ning Guo, Quanzheng Li

专题命中 多模态训练与对齐 :multi-modal(title,abstract);multimodal(abstract);分类 cs.CV

Comments Zhe Guo and Xiang Li contribute equally to this work

Journal ref 2018 IEEE 15th International Symposium on Biomedical Imaging (ISBI 2018), Washington, DC, 2018, pp. 903-907

详情

展开后加载摘要…

URL PDF HTML 收藏
1611.04503 2017-07-25 cs.CL cs.AI cs.CV cs.MM 83%

Zero-resource Machine Translation by Multimodal Encoder-decoder Network with Multimedia Pivot

Hideki Nakayama, Noriki Nishida

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Some error corrections in Sect.2.2 and Table 5, Machine Translation, 2017

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.22709 2026-07-28 cs.CV cs.AI cs.CL 新提交 83%

RMS@CC-MMD 2026: Multimodal Misogyny Detection via Geometric Interaction and Multi-View Consensus

RMS@CC-MMD 2026:通过几何交互和多视图共识进行多模态厌女症检测

Md. Ajwad Hossain

机构 * Chittagong University of Engineering & Technology (CUET)(吉大港工程技术大学)

专题命中 多模态训练与对齐 :multimodal(title,comments);cross-modal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 研究针对多模态厌女症检测问题,提出GeoMVC方法,通过几何交互层建模跨模态对齐,用多视图共识策略减轻分布偏移,在相关挑战中取得一定排名成绩,凸显特定文化背景下建模的挑战。

Comments Accepted for the CC-MMD Grand Challenge at the 28th ACM International Conference on Multimodal Interaction (ICMI 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.17372 2024-10-08 cs.IR 83%

An Empirical Study of Training ID-Agnostic Multi-modal Sequential Recommenders

Youhua Li, Hanwen Du, Yongxin Ni, Yuanqi He, Junchen Fu, Xiangyan Liu, Qi Guo

专题命中 多模态训练与对齐 :multi-modal(title,abstract);multimodal(abstract)

Comments An Empirical Study of Training ID-Agnostic Multi-modal Sequential Recommenders

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.00102 2023-04-10 cs.CV cs.AI cs.MM 83%

Dynamic Multimodal Fusion

Zihui Xue, Radu Marculescu

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI、cs.MM;multi-modal(comments)

Comments Accepted by 6th Multi-Modal Learning and Applications Workshop (MULA), CVPR 2023. Code available at: https://github.com/zihuixue/DynMM

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.17444 2022-12-06 cs.LG 83%

Multimodal Information Bottleneck: Learning Minimal Sufficient Unimodal and Multimodal Representations

Sijie Mai, Ying Zeng, Haifeng Hu

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)

Comments This paper is accepted by IEEE Transactions on Multimedia. This version addresses some mistakes and typos in the original paper. The appendix is available at https://github.com/TmacMai/Multimodal-Information-Bottleneck/blob/main/appendix.pdf

详情

展开后加载摘要…

URL PDF HTML 收藏
1808.07275 2018-08-23 cs.AI cs.CV cs.MM 83%

CentralNet: a Multilayer Approach for Multimodal Fusion

Valentin Vielzeuf, Alexis Lechervy, Stéphane Pateux, Frédéric Jurie

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI、cs.MM

Journal ref European Conference on Computer Vision Workshops: Multimodal Learning and Applications, Sep 2018, Munich, Germany. https://mula2018.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.18132 2026-08-20 cs.CL cs.SD eess.AS 新提交 82%

Alignment Is All You Need: Instruction-Free Training for General Audio-Language Models

对齐即全部所需:通用音频-语言模型的无指令训练

Xuanru Zhou, Yiwen Shao, Jiahong Li, Dong Yu

机构 * Zhejiang University(浙江大学) Tencent Hunyuan(腾讯混元)

专题命中 多模态训练与对齐 :MLLM(abstract,abstract_cn);multimodal(abstract);cross-modal(abstract);分类 cs.CL、eess.AS

AI总结 该研究提出仅对齐的无指令大音频-语言模型LALM,仅训练轻量投影器,在多模态数据集上用更少数据达到或优于基线,证明仅靠对齐即可构建有竞争力的多模态大语言模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11171 2026-06-19 cs.CV cs.AI 版本更新 82%

TerraMind: Large-Scale Generative Multimodality for Earth Observation

TerraMind:面向地球观测的大规模生成式多模态模型

Johannes Jakubik, Felix Yang, Benedikt Blumenstiel, Erik Scheurer, Rocco Sedona, Stefano Maurogiovanni, Jente Bosmans, Nikolaos Dionelis, Valerio Marsocci, Niklas Kopp, Rahul Ramachandran, Paolo Fraccaro, Thomas Brunschwiler, Gabriele Cavallaro, Juan Bernabe-Moreno, Nicolas Longépé

机构 * IBM Research – Europe(IBM欧洲研究院) ETH Zurich(苏黎世联邦理工学院) Forschungszentrum Jülich(尤利希研究中心) European Space Agency(欧洲航天局) Φ \Phi -Lab(Φ实验室) NASA IMPACT University of Iceland(爱沙尼亚大学)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);any-to-any(abstract);multimodal foundation model(abstract)

AI总结 提出首个任意到任意生成式多模态基础模型TerraMind,通过双尺度表示(token级和像素级)预训练,实现零样本/少样本应用,并引入“模态思考”能力,在PANGAEA等基准上达到领先性能。

Comments Accepted at ICCV'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18681 2025-07-04 cs.CL cs.AI 82%

Commander-GPT: Fully Unleashing the Sarcasm Detection Capability of Multi-Modal Large Language Models

Yazhou Zhang, Chunwang Zou, Bo Wang, Jing Qin

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CL、cs.AI;multimodal(comments)

Comments Our original goal was to use Commander-GPT: Dividing and Routing for Multimodal Sarcasm Detection (arXiv:2506.19420) to replace Commander-GPT: Fully Unleashing the Sarcasm Detection Capability of Multi-Modal Large Language Models (arXiv:2503.18681). Due to various reasons, both versions were released, so we would like to withdraw the latter

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.19597 2026-08-21 cs.LG stat.ML 82%

The Geometric Mechanics of Contrastive Representation Learning: Alignment Potentials, Entropic Dispersion, and Cross-modal Divergence

对比表示学习的几何力学:对齐势、熵分散和跨模态散度

Yichao Cai, Zhen Zhang, Yuhang Liu, Javen Qinfeng Shi

机构 * University of Science and Technology of China(中国科学技术大学)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract)

AI总结 本文通过测度论框架,在大批量极限下证明InfoNCE目标与确定性能量景观的等价性,揭示单模态与对称多模态之间的几何分岔,并指出跨模态散度项导致模态间隙。

Comments Accepted at ICML 2026; v7: Exposition and notation refined; results unchanged

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.16201 2026-08-18 cs.LG 新提交 82%

Multi-Granularity Sentiment Integration for LLM-Based Multimodal Sentiment Analysis

面向基于大语言模型(LLM)的多模态情感分析的多粒度情感集成

Shanshan Lin, Yuesheng Wu, Chao Chen, Yizhe Yang, Zhihao Chen, Zexian Yang, Xiangwen Liao

机构 * Fuzhou University(福州大学) Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳)) Jiangxia University(江夏大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)

AI总结 该研究提出MGSI多粒度情感集成框架,通过多尺度编码、文本引导对齐等优化,提升基于LLM的多模态情感分析性能,在四个公开基准上效果优于冻结LLM基线。

Comments Accepted to NLPCC 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.09572 2026-08-11 cs.LG 新提交 82%

Hyperbolic Multimodal Continual Learning

双曲多模态持续学习

Jiahong Liu, Ming Shen, Xiaohao Liu, Rex Ying, Menglin Yang, Tat-Seng Chua, Irwin King

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)

AI总结 本研究针对双曲多模态持续学习的遗忘问题,建立理论基础推导了保留几何结构的持续学习框架,经实验验证其有效性。

Comments ICML 2026. 33 pages, 11 figures. Code: ICML" target="_blank" rel="noopener">https://github.com/HUBERILT/HMCL_ICML

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.07895 2026-08-11 cs.RO cs.LG 新提交 82%

Auditing Instruction-Trajectory Mismatches in Multimodal Robot Demonstrations

多模态机器人演示中指令-轨迹不匹配的审计

Simon Holk, Ryosuke Takanami, Tatsuya Matsushima, Yusuke Iwasawa, Yutaka Matsuo, Yueh-Hua Wu, Kei Ota

机构 * AI Robot Association (AIRoA)(人工智能机器人协会(AIRoA)) The University of Tokyo(东京大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract)

AI总结 针对多模态机器人演示中指令-轨迹不匹配问题,提出无需训练的MMPF审计框架,在LIBERO基准及真实机器人数据上实现最优ITM检测与标签修正,可提升下游策略学习性能并展示过滤演示的权衡。

Comments Accepted for publication in IEEE Robotics and Automation Letters (RA-L). 8 pages, 3 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06211 2026-08-11 cs.CL cs.AI eess.AS 版本更新 82%

LF${}^{2}$AR: Accounting for Layerwise Dynamics to Improve Multimodal Adaptation of Language Models

LF²AR:考虑分层动态以改进语言模型的多模态适配

Santiago Cuervo, Adel Moumen, Yanis Labrak, Sameer Khurana, Antoine Laurent, Mickael Rouvier, Phil Woodland, Ricard Marxer

机构 * Université de Toulon, Aix-Marseille Université, CNRS, LIS, France(法国图卢兹大学、马赛大学、CNRS、LIS) Department of Engineering, University of Cambridge, UK(剑桥大学工程系) Mitsubishi Electric Research Laboratories (MERL), Cambridge, MA, USA(三菱电机研究实验室(MERL)) LIA, Avignon Université, France(法国阿维尼翁大学LIA) LIUM, Le Mans Université, France(法国勒芒大学LIUM) Zenidoc, Marseille, France(法国马赛Zenidoc)

专题命中 多模态训练与对齐 :multimodal(title);cross-modal(abstract);分类 cs.CL、cs.AI、eess.AS

AI总结 本研究提出LF²AR架构,通过分层抽象-细化动态设计适配机制,在文本转图像、语音模态上提升语言模型性能,支持1.9倍生成加速。

Comments Published as a conference paper at COLM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏