arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6918 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6918 篇

1510.03519 2016-07-04 cs.CL 79%

Bridge Correlational Neural Networks for Multilingual Multimodal Representation Learning

Janarthanan Rajendran, Mitesh M. Khapra, Sarath Chandar, Balaraman Ravindran

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

Comments Published at NAACL-HLT 2016

详情

展开后加载摘要…

URL PDF HTML 收藏
1604.03443 2016-04-13 cs.CV 79%

Multi-modal Fusion for Diabetes Mellitus and Impaired Glucose Regulation Detection

Jinxing Li, David Zhang, Yongcheng Li, Jian Wu

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments 9 pages, 8 figures, 30 conference

详情

展开后加载摘要…

URL PDF HTML 收藏
1502.01094 2016-01-20 stat.ML cs.CV cs.LG 79%

Multimodal Task-Driven Dictionary Learning for Image Classification

Soheil Bahrampour, Nasser M. Nasrabadi, Asok Ray, W. Kenneth Jenkins

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments To appear at IEEE Transactions on Image Processing

详情

展开后加载摘要…

URL PDF HTML 收藏
1502.07432 2015-11-19 cs.CV physics.data-an 79%

Coercive Region-level Registration for Multi-modal Images

Yu-Hui Chen, Dennis Wei, Gregory Newstadt, Jeffrey Simmons, Alfred Hero

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments This work has been accepted to International Conference on Image Processing (ICIP) 2015

详情

展开后加载摘要…

URL PDF HTML 收藏
1501.00102 2015-07-21 cs.CV cs.HC cs.LG 79%

ModDrop: adaptive multi-modal gesture recognition

Natalia Neverova, Christian Wolf, Graham W. Taylor, Florian Nebout

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments 14 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1411.3229 2015-01-14 cs.CV 79%

Multi-modal Image Registration for Correlative Microscopy

Tian Cao, Christopher Zach, Shannon Modla, Debbie Powell, Kirk Czymmek, Marc Niethammer

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments 24 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
1011.6220 2010-11-30 cs.AI 79%

Multimodal Biometric Systems - Study to Improve Accuracy and Performance

K. Sasidhar, Vijaya L Kakulapati, Kolikipogu Ramakrishna, K. KailasaRao

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments 8 pages,5 figures, published in International Journal of Computer Science & Engineering Survey (IJCSES) Vol.1, No.2, November 2010

详情

展开后加载摘要…

URL PDF HTML 收藏
0909.4280 2009-12-01 cs.CL 79%

Towards Multimodal Content Representation

Harry Bunt, Laurent Romary

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

Comments Colloque avec actes et comité de lecture. internationale

Journal ref LREC Workshop on International Standards of Terminology and Language Resources Management, Las Palams : Spain (2002)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.15127 2025-11-19 cs.LG 79%

PRIMUS: Pretraining IMU Encoders with Multimodal Self-Supervision

Arnav M. Das, Chi Ian Tang, Fahim Kawsar, Mohammad Malekzadeh

机构 * Nokia Bell Labs Cambridge, UK(诺基亚贝尔实验室(剑桥,英国)) University of Washington, USA(华盛顿大学(美国)) University of Glasgow, UK(格拉斯哥大学(英国))

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments Presented at ICASSP 2025. Also presented under the title "PRIMUS: Pretraining IMU Encoders with Multimodal and Self-Supervised Learning" at NeurIPS 2024 TSALM Workshop (Time Series in the Age of Large Models)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13637 2025-11-18 cs.LG 79%

Towards Multimodal Representation Learning in Paediatric Kidney Disease

Ana Durica, John Booth, Ivana Drobnjak

机构 * Institute of Health Informatics(健康信息学研究所) University College London(伦敦大学学院) Data Research, Innovation and Virtual Environments Unit(数据研究、创新与虚拟环境单位) Great Ormond Street Hospital(格雷特奥蒙德医院) Department of Computer Science(计算机科学系)

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments 4 pages, 3 figures. EurIPS 2025 Multimodal Representation Learning for Healthcare (MMRL4H) workshop paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21971 2025-11-05 cs.LG 79%

GRAM-DTI: adaptive multimodal representation learning for drug target interaction prediction

Feng Jiang, Amina Mollaysa, Hehuan Ma, Tommaso Mansi, Junzhou Huang, Mangal Prakash, Rui Liao

机构 * University of Texas at Arlington(德克萨斯理工大学) Johnson & Johnson Innovative Medicine(强生创新医学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(journal_ref)

Journal ref NeurIPS 2025 2nd Workshop on Multi-modal Foundation Models and Large Language Models for Life Sciences

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16026 2025-11-03 cs.LG stat.AP 79%

A tutorial on discovering and quantifying the effect of latent causal sources of multimodal EHR data

Marco Barbero-Mota, Eric V. Strobl, John M. Still, William W. Stead, Thomas A. Lasko

机构 * Department of Biomedical Informatics Vanderbilt University Medical Center(生物医学信息学系范德堡大学医学中心) Department of Biomedical Informatics University of Pittsburgh(生物医学信息学系匹兹堡大学) Departments of Medicine & Biomedical Informatics Vanderbilt University Medical Center(医学与生物医学信息学系范德堡大学医学中心) Departments of Biomedical Informatics & Computer Science Vanderbilt University Medical Center & Vanderbilt University(生物医学信息学与计算机科学系范德堡大学医学中心及范德堡大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments Accepted at the 1st Multimodal Representation Learning for Healthcare EurIPS 2025 Workshop (https://multimodal-rep-learning-for-health.github.io/)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05993 2025-10-24 cs.IR 79%

Efficient Multimodal Streaming Recommendation via Expandable Side Mixture-of-Experts

Yunke Qu, Liang Qu, Tong Chen, Quoc Viet Hung Nguyen, Hongzhi Yin

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments Accepted to CIKM 2025. Code is available at https://github.com/qykcq/Efficient-Multimodal-Streaming-Recommendation-via-Expandable-Side-Mixture-of-Experts

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08936 2025-06-11 cs.LG 79%

BioLangFusion: Multimodal Fusion of DNA, mRNA, and Protein Language Models

Amina Mollaysa, Artem Moskale, Pushpak Pati, Tommaso Mansi, Mangal Prakash, Rui Liao

专题命中 多模态训练与对齐 :multimodal(title);cross-modal(abstract);multi-modal(comments)

Comments Proceedings of ICML 2025 Workshop on Multi-modal Foundation Proceedings of ICML 2025 Workshop on Multi-modal Foundation Proceedings of ICML 2025 Workshop on Multi-modal Foundation Models and Large Language Models for Life Sciences

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.01215 2024-12-03 cs.LG 79%

EsurvFusion: An evidential multimodal survival fusion model based on Gaussian random fuzzy numbers

Ling Huang, Yucheng Xing, Qika Lin, Su Ruan, Mengling Feng

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments Multimodal survival analysis, Epistemic random fuzzy sets theory, Uncertainty

详情

展开后加载摘要…

URL PDF HTML 收藏
1903.07303 2019-03-19 cs.LG stat.ML 79%

M$^2$VAE - Derivation of a Multi-Modal Variational Autoencoder Objective from the Marginal Joint Log-Likelihood

Timo Korthals

专题命中 多模态训练与对齐 :multi-modal(title,abstract)

Comments Appendix for the IEEE FUSION 2019 submission on multi-modal variational Autoencoders for sensor fusion

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.19726 2026-08-21 cs.CL cs.CV cs.LG 新提交 79%

Projector Is All You Train

仅训练投影器就足够

Nyx Iskandar, Saathvik Selvan, Slater Victoroff

机构 * Ramen VR(拉面VR公司) University of California, Berkeley(加州大学伯克利分校)

专题命中 多模态训练与对齐 :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV、cs.CL

AI总结 该研究探究多模态大语言模型适配新模态是否需微调主干,经实验发现仅训练投影器即可实现强多模态性能,还能避免联合训练导致的语言模型能力漂移,且训练样本吞吐量约为联合训练的两倍,通过多类基准验证了结论。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.15821 2026-08-12 cs.CL cs.AI cs.LG 版本更新 79%

The Truth Stays in the Family: Enhancing Contextual Grounding via Inherited Truthful Heads in Model Lineages

真相留在家族中:通过模型谱系中继承的真相头增强上下文基础

Miso Choi, Seonga Choi, Mincheol Kwon, Woosung Joung, Jinkyu Kim, Jungbeom Lee

机构 * University of Science and Technology of China(中国科学技术大学)

专题命中 多模态训练与对齐 :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CL、cs.AI

AI总结 研究发现基础LLM与下游变体间存在上下文真相分数的强继承性,提出TruthProbe软门控策略放大真相头以提升上下文真实性并减少多模态幻觉。

Comments Accepted at ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.05131 2026-08-07 cs.CV cs.AI 版本更新 79%

OPD-V: Visual On-Policy Self-Distillation with Modality Balance

OPD-V:结合模态平衡的视觉在线策略自蒸馏

Aniri, Jinhe Bi, Peng Liao, Zengjie Jin, Volker Tresp, Fei Shen, Yunpu Ma, Tat-Seng Chua

机构 * National University of Singapore(新加坡国立大学) Ludwig Maximilian University of Munich(慕尼黑大学) Munich Center for Machine Learning(慕尼黑机器学习中心) Sun Yat-sen University(中山大学)

专题命中 多模态训练与对齐 :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV、cs.AI

AI总结 该研究针对多模态大语言模型的模态不平衡问题,提出视觉在线策略自蒸馏范式OPD-V,通过正、负教师模型实现模态平衡,在多基准与骨干上提升推理性能并降低训练成本。

Comments Corrected the uploaded manuscript. Project Page:https://github.com/aniri15/OPD-V

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.23445 2026-07-28 cs.CV cs.CL 新提交 79%

Omni-Prune: Query-Aware Unified Token Pruning for Efficient Omnimodal Large Language Models

Omni-Prune:用于高效全模态大语言模型的查询感知统一令牌剪枝

Yiming Zhong, Chang Nie, Caifeng Shan

机构 * Nanjing University(南京大学)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);audio-visual(abstract);分类 cs.CV、cs.CL

AI总结 研究针对全模态大语言模型推理时音频-视频令牌序列长、预填充延迟高和GPU内存使用量大的问题,提出无需训练的查询感知Omni-Prune框架,联合去除冗余并保留跨模态证据,实验证明其性能优于基线方法,能加速预填充并减少内存。

Comments 14 pages, 7 figures. Code: https://github.com/kimberlyii/Omni-Prune

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.04079 2026-07-07 cs.CV cs.AI cs.LG 新提交 79%

Seeing Once is Enough? Online Geometry-Aware Token Pruning for 3D Question Answering

看一次就够了?用于3D问答的在线几何感知令牌剪枝

Ruei-Chi Lai, Bolivar Solarte, Chin-Hsuan Wu, Yi-Hsuan Tsai, Min Sun

机构 * National Tsing Hua University(国立清华大学) Industrial Technology Research Institute ITRI(工业技术研究院) University of Toronto(多伦多大学) Atmanity Inc(Atmanity公司)

专题命中 多模态训练与对齐 :MLLM(abstract,abstract_cn);multi-modal(abstract);分类 cs.CV、cs.AI

AI总结 针对3D问答中多模态大语言模型推理成本高的问题,提出在线令牌剪枝方法,利用深度和相机姿态投影到体素空间,识别重叠区域并剪枝冗余令牌,减少令牌使用,提升效率和性能。

Comments published at ICLR 2026 Workshop on Efficient Spatial Reasoning

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.22138 2026-06-23 cs.CL cs.AI cs.LG q-bio.BM 新提交 79%

BioMatrix: Towards a Comprehensive Biological Foundation Model Spanning the Modality Matrix of Sequences, Structures, and Language

BioMatrix:迈向覆盖序列、结构和语言模态矩阵的综合性生物基础模型

Qizhi Pei, Zhimeng Zhou, Yi Duan, Yiyang Zhao, Wei Li, Han Guo, Liang He, Chengping Li, Chang-Yu Hsieh, Conghui He, Rui Yan, Lijun Wu

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学高瓴人工智能学院) OpenDataLab, Shanghai Artificial Intelligence Laboratory(上海人工智能实验室 OpenDataLab) Zhejiang University(浙江大学) Shanghai Innovation Institute(上海创新研究院) East China Normal University(华东师范大学) Zhongguancun Academy(中关村学院) School of Artificial Intelligence, Wuhan University(武汉大学人工智能学院)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);multimodal foundation model(abstract);分类 cs.CL、cs.AI

AI总结 提出首个原生多模态生物基础模型BioMatrix,通过统一分词方案将分子序列、结构、蛋白质序列、结构和自然语言映射到共享离散标记空间,在单一解码器架构下实现所有模态的统一生成,在80项任务中77项达到最优或竞争力水平。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.28422 2026-05-28 cs.CV cs.AI 79%

VITAL: Visual-Semantic Dual Supervision for Enhanced and Interpretable Latent Reasoning in Medical MLLMs

VITAL: 视觉-语义双重监督增强可解释的医学多模态大语言模型潜在推理

Qiaoru Li, Shaotian Liang, Jintao Chen, Haoran Sun, Yuxiang Cai, Jianwei Yin, Yankai Jiang

机构 * Zhejiang University(浙江大学) Shanghai AI Laboratory(上海人工智能实验室) Tencent(腾讯) Ningbo Global Innovation Center, Zhejiang University(宁波全球创新中心,浙江大学) Zhejiang Key Laboratory of Digital-Intelligence Service Technology(浙江省数字智能服务技术重点实验室)

专题命中 多模态训练与对齐 :MLLM(summary_cn,abstract_cn);分类 cs.CV、cs.AI

AI总结 提出VITAL框架,通过视觉-语义双重监督(文本解码器重构推理链、视觉投影器回归ROI特征)实现医学MLLM的可解释潜在推理,在7个基准上达到SOTA。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.22036 2026-05-22 cs.CV cs.AI 79%

GA-VLN: Geometry-Aware BEV Representation for Efficient Vision-Language Navigation

GA-VLN: 用于高效视觉-语言导航的几何感知鸟瞰图表示

Jiahao Yang, Zihan Wang, Xiangyang Li, Xing Zhu, Yujun Shen, Yinghao Xu, Shuqiang Jiang

机构 * State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences(人工智能安全国家重点实验室,计算技术研究所,中国科学院) University of Chinese Academy of Sciences(中国科学院大学) Robbyant School of Computing, National University of Singapore(新加坡国立大学计算机学院) The Hong Kong University of Science and Technology(香港科技大学)

专题命中 多模态训练与对齐 :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出GA-VLN框架,通过引入几何感知的鸟瞰图表示(GA-BEV),整合显式和隐式几何信息,提升视觉-语言导航的效率和性能,实验表明其在仅使用导航数据的情况下取得了最先进的结果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.01402 2026-05-12 cs.CL cs.CV cs.LG 79%

Injecting Distributional Awareness into MLLMs via Reinforcement Learning for Deep Imbalanced Regression

通过强化学习为深度不平衡回归注入分布意识

Yao Du, Shanshan Song, Xiaomeng Li

机构 * The Hong Kong University of Science(香港科学与技术大学)

专题命中 多模态训练与对齐 :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV、cs.CL

AI总结 本文提出基于组相对策略优化的分布感知强化学习框架,通过一致性相关系数奖励实现跨样本关系监督,提升长尾回归任务的分布对齐能力,在中等和少样本场景下表现优异。

Comments Accepted by ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.07825 2026-05-11 cs.MM cs.CV 79%

Anisotropic Modality Align

各向异性模态对齐

Xiaomin Yu, Yijiang Li, Yuhui Zhang, Hanzhen Zhao, Yue Yang, Hao Tang, Yue Song, Xiaobin Hu, Chengwei Qin, Shuicheng Yan, Hui Xiong

机构 * HKUST(GZ)(香港科技大学(广州)) NUS(国立新加坡大学) UCSD(加州大学圣地亚哥分校) Stanford(斯坦福大学) PKU(北京大学) THU(清华大学)

专题命中 多模态训练与对齐 :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV、cs.MM

AI总结 本文研究了多模态模型中模态间转换的可行性,提出各向异性模态对齐方法,通过几何修正框架提升单模态数据的多模态训练效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02767 2026-03-10 cs.CV cs.AI 79%

ITO: Images and Texts as One via Synergizing Multiple Alignment and Training-Time Fusion

通过协同多模态对齐和训练时融合实现图像与文本一体化:ITO

Hanpeng Liu, Yaqian Li, Zidan Wang, Shuoxi Zhang, Zonglin Zhao, Zihao Bo, Rinyoichi Takezoe, Kaiwen Long, Kun He

机构 * School of Computer Science(计算机科学学院) Huazhong University of Science and Technology(华中科技大学) Li Auto Inc.(力汽车公司) Institute of AI for Industries, Chinese Academy of Sciences(产业人工智能研究院,中国科学院)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);image-text(abstract);分类 cs.CV、cs.AI

AI总结 ITO通过协同多模态对齐与训练时融合机制,提升图像与文本表示的一致性,有效解决模态间结构化交互问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10161 2026-02-12 cs.CR cs.AI cs.CL 79%

Omni-Safety under Cross-Modality Conflict: Vulnerabilities, Dynamics Mechanisms and Efficient Alignment

跨模态冲突下的全方位安全:漏洞、动态机制和高效对齐

Kun Wang, Zherui Li, Zhenhong Zhou, Yitong Zhang, Yan Mi, Kun Yang, Yiming Zhang, Junhao Dong, Zhongxiang Sun, Qiankun Li, Yang Liu

机构 * Nanyang Technological University(南洋理工大学) Beijing University of Posts and Telecommunications(北京邮电大学) Tsinghua University(清华大学) Fudan University(复旦大学) University of Science and Technology of China(中国科学技术大学) Renmin University of China(中国人民大学)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);omni-modal(abstract);分类 cs.CL、cs.AI

AI总结 本文提出OmniSteer方法,通过提取黄金拒绝向量和轻量级适配器提升多模态模型的安全性与通用能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03019 2025-11-06 cs.CV cs.AI 79%

SLIP: Structural-aware Language-Image Pretraining for Vision-Language Alignment

Wenbo Lu

机构 * Department of Data Science(数据科学系) New York University Shanghai(纽约大学上海分校)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);image-text(abstract);分类 cs.CV、cs.AI

Comments Capstone Paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01357 2025-11-04 cs.CV cs.AI 79%

CMI-MTL: Cross-Mamba interaction based multi-task learning for medical visual question answering

Qiangguo Jin, Xianyao Zheng, Hui Cui, Changming Sun, Yuqi Fang, Cong Cong, Ran Su, Leyi Wei, Ping Xuan, Junbo Wang

机构 * School of Software, Northwestern Polytechnical University, Shaanxi, China(西北工业大学软件学院) Yangtze River Delta Research Institute of Northwestern Polytechnical University, Taicang, China(西北工业大学长江三角研究 institute) Department of Computer Science and Information Technology, La Trobe University, Melbourne, Australia(拉筹伯大学计算机科学与信息技术系) CSIRO Data61, Sydney, Australia(CSIRO Data61) School of Intelligence Science and Technology, Nanjing University, Suzhou, China(南京大学智能科学与技术学院) Australian Institute of Health Innovation (AIHI), Macquarie University, Australia(麦考瑞大学健康创新研究所) School of Computer Software, College of Intelligence and Computing, Tianjin University, Tianjin, China(天津大学计算机软件学院) Centre for Artificial Intelligence driven Drug Discovery, Faculty of Applied Science, Macao Polytechnic University, Macao Special Administrative Region of China(澳门理工学院人工智能驱动药物发现中心) Department of Computer Science, School of Engineering, Shantou University, Guangdong, China(汕头大学计算机科学系)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);image-text(abstract);分类 cs.CV、cs.AI

Comments The paper has been accepted by the 33rd Pacific Conference on Computer Graphics and Applications (Pacific Graphics 2025)

Journal ref PG2025 Conference Papers, Posters, and Demos, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏