arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6918 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6918 篇

1706.00153 2017-06-27 cs.MM cs.CV cs.LG 81%

Cross-modal Common Representation Learning by Hybrid Transfer Network

Xin Huang, Yuxin Peng, Mingkuan Yuan

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV、cs.MM

Comments To appear in the proceedings of 26th International Joint Conference on Artificial Intelligence (IJCAI), Melbourne, Australia, Aug. 19-25, 2017. 8 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
cs/0703091 2016-08-14 cs.AI cs.MM 81%

Multimodal Meaning Representation for Generic Dialogue Systems Architectures

Frédéric Landragin, Alexandre Denis, Annalisa Ricci, Laurent Romary

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI、cs.MM

Journal ref Proceedings of the Fourth International Conference on Language Resources and Evaluation (LREC 2004) (2004) 521-524

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.22158 2026-03-24 cs.LG cs.AI 80%

Multimodal Survival Analysis with Locally Deployable Large Language Models

多模态生存分析与可本地部署的大语言模型

Moritz Gögl, Christopher Yau

机构 * University of Oxford(牛津大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI;multi-modal(comments)

AI总结 本文提出利用可本地部署的大语言模型进行多模态生存分析,结合临床文本、表格数据和基因组数据,通过教师-学生蒸馏和原理化的多模态融合,实现校准的生存概率估计和简洁的诊断文本生成,优于标准基线并在隐私和准确性方面表现更优。

Comments NeurIPS 2025 Workshop on Multi-modal Foundation Models and Large Language Models for Life Sciences

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.18652 2026-03-17 cs.CL 80%

PolyFrame at MWE-2026 AdMIRe 2: When Words Are Not Enough: Multimodal Idiom Disambiguation

PolyFrame在MWE-2026 AdMIRe 2:当词语不够时:多模态成语消歧

Nina Hosseini-Kivanani

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

AI总结 本文提出PolyFrame系统,通过统一管道处理图像+文本排名和纯文本描述排名任务,利用轻量模块提升多语言成语消歧性能,无需微调大模型。

Comments Accepted at AdMIRe 2 shared task (Advancing Multimodal Idiomaticity Representation) colocated with 22nd Workshop on Multiword Expressions (MWE 2026) @EACL2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.03622 2026-02-04 cs.CV physics.med-ph 80%

Quasi-multimodal-based pathophysiological feature learning for retinal disease diagnosis

基于准多模态的病理特征学习用于视网膜疾病诊断

Lu Zhang, Huizhen Yu, Zuowei Wang, Fu Gui, Yatu Guo, Wei Zhang, Mengyu Jia

机构 * Tianjin University(天津大学) Tianjin Key Laboratory of Ophthalmology and Visual Science(天津眼科学与视觉科学重点实验室) Tianjin Eye Institute(天津眼科研究院) Tianjin Eye Hospital(天津眼科医院) Clinical College of Ophthalmology, Tianjin Medical University(天津医科大学临床医学院) Department of Ophthalmology, The Second Affiliated Hospital of Nanchang University(南昌大学第二附属医院眼科部) Nankai University Affiliated Eye Hospital(南开大学附属眼科医院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出基于准多模态的视网膜疾病诊断方法,通过多模态数据合成与融合提升分类和分级的准确性。

Journal ref Zhang, L., Yu, H., Wang, Z., Gui, F., Guo, Y., Zhang, W., Jia, M., 2026. Quasi-multimodal-based pathophysiological feature learning for retinal disease diagnosis. Medical Image Analysis 109, 103886

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09470 2026-01-15 physics.ed-ph cs.AI 80%

Personalized Multimodal Feedback Using Multiple External Representations: Strategy Profiles and Learning in High School Physics

基于多种外部表征的个性化反馈:策略配置与高中物理学习中的学习

Natalia Revenga-Lozano, Karina E. Avila, Steffen Steinert, Matthias Schweinberger, Clara E. Gómez-Pérez, Jochen Kuhn, Stefan Küchemann

机构 * Chair of Physics Education, Faculty of Physics, Ludwig-Maximilians-Universität München (LMU Munich)(物理教育系主任,物理学院,慕尼黑路易斯-马克西姆利安大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

AI总结 本文研究了多种外部表征与个性化反馈在高中物理学习中的整合效果,发现详细多表征反馈对学习成绩有积极影响,且学习者根据表征能力选择不同反馈策略。

Comments Keywords: Adaptive Feedback, Multimodal Learning, Multiple External Representations, Physics Education, Science Education, Representational Competences, Intelligent Tutoring Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15986 2025-11-25 cs.CV cs.CY cs.LG 80%

Fairness in Multi-modal Medical Diagnosis with Demonstration Selection

多模态医学诊断中的公平性与演示选择

Dawei Li, Zijian Gu, Peng Wang, Chuhan Song, Zhen Tan, Mohan Zhang, Tianlong Chen, Yu Tian, Song Wang

机构 * Arizona State University(亚利桑那州立大学) University of Rochester(罗切斯特大学) University of Virginia(弗吉尼亚大学) UCL(伦敦大学学院) University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校) University of Central Florida(佛罗里达中央大学)

专题命中 多模态训练与对齐 :multi-modal(title,comments);multimodal(abstract);分类 cs.CV

AI总结 本文提出FADS方法,通过基于聚类的采样提升多模态医学影像诊断的公平性,减少性别、种族和族裔相关差异,同时保持高准确性。

Comments 10 pages (including 2 pages of references), 4 figures. This work explores fairness in multi-modal medical image reasoning using in-context learning

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.05177 2025-10-29 cs.CV 80%

Long-VITA: Scaling Large Multi-modal Models to 1 Million Tokens with Leading Short-Context Accuracy

Yunhang Shen, Chaoyou Fu, Shaoqi Dong, Xiong Wang, Yi-Fan Zhang, Peixian Chen, Mengdan Zhang, Haoyu Cao, Ke Li, Shaohui Lin, Xiawu Zheng, Yan Zhang, Yiyi Zhou, Ran He, Caifeng Shan, Rongrong Ji, Xing Sun

机构 * Tencent Youtu Lab(腾讯云图实验室) Nanjing University(南京大学) East China Normal University(华东师范大学) Xiamen University(厦门大学) CASIA(中国科学院自动化研究所)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV;MLLM(comments)

Comments https://github.com/VITA-MLLM/Long-VITA

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.10353 2025-04-21 cs.CV cs.CR cs.LG 80%

Robust image classification with multi-modal large language models

Francesco Villani, Igor Maljkovic, Dario Lazzaro, Angelo Sotgiu, Antonio Emanuele Cinà, Fabio Roli

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV;multimodal(comments)

Comments Paper accepted at Pattern Recognition Letters journal Keywords: adversarial examples, rejection defense, multimodal-informed systems, machine learning security

Journal ref Pattern Recognition Letters 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.14504 2025-03-25 cs.CV 80%

Aligning Multimodal LLM with Human Preference: A Survey

Tao Yu, Yi-Fan Zhang, Chaoyou Fu, Junkang Wu, Jinda Lu, Kun Wang, Xingyu Lu, Yunhang Shen, Guibin Zhang, Dingjie Song, Yibo Yan, Tianlong Xu, Qingsong Wen, Zhang Zhang, Yan Huang, Liang Wang, Tieniu Tan

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Project page: https://github.com/BradyFU/Awesome-Multimodal-Large-Language-Models/tree/Alignment

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.02884 2024-06-12 cs.CV 80%

Vision+X: A Survey on Multimodal Learning in the Light of Data

Ye Zhu, Yu Wu, Nicu Sebe, Yan Yan

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Survey paper on multimodal learning and generation, to appear at IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI)

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.13765 2024-05-21 cs.AI 80%

Towards ethical multimodal systems

Alexis Roger, Esma Aïmeur, Irina Rish

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments 5 pages, multimodal ethical dataset building, accepted in the NeurIPS 2023 MP2 workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.04125 2023-11-01 cs.LG cs.CL cs.HC 80%

Multimodal Fusion Interactions: A Study of Human and Automatic Quantification

Paul Pu Liang, Yun Cheng, Ruslan Salakhutdinov, Louis-Philippe Morency

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

Comments International Conference on Multimodal Interaction (ICMI '23), Code available at: https://github.com/pliang279/PID. arXiv admin note: text overlap with arXiv:2302.12247

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.02327 2021-11-04 cs.HC cs.CV 80%

ML-PersRef: A Machine Learning-based Personalized Multimodal Fusion Approach for Referencing Outside Objects From a Moving Vehicle

Amr Gomaa, Guillermo Reyes, Michael Feld

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Journal ref In Proceedings of the 2021 International Conference on Multimodal Interaction, pp. 318-327. 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2007.02038 2020-07-07 cs.CL 80%

Low Rank Fusion based Transformers for Multimodal Sequences

Saurav Sahay, Eda Okur, Shachi H Kumar, Lama Nachman

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

Comments ACL 2020 workshop on Second Grand Challenge and Workshop on Multimodal Language

详情

展开后加载摘要…

URL PDF HTML 收藏
1609.05281 2016-09-20 cs.CV 80%

GeThR-Net: A Generalized Temporally Hybrid Recurrent Neural Network for Multimodal Information Fusion

Ankit Gandhi, Arjun Sharma, Arijit Biswas, Om Deshmukh

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV;audio-visual(comments)

Comments To appear in ECCV workshop on Computer Vision for Audio-Visual Media, 2016

详情

展开后加载摘要…

URL PDF HTML 收藏
0909.2373 2009-12-01 cs.CV cs.CR 80%

An Efficient Secure Multimodal Biometric Fusion Using Palmprint and Face Image

M. Nageshkumar, P. K. Mahesh, M. N. S. Swamy

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments International Journal of Computer Science Issues (IJCSI), Volume 1, pp49-53, August 2009

Journal ref M.Nageshkumar,P.K.Mahesh and M.N.S.Swamy, "An Efficient Secure Multimodal Biometric Fusion Using Palmprint and Face Image", International Journal of Computer Science Issues (IJCSI), Volume 1, pp49-53, August 2009

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11693 2025-10-14 cs.CL cs.AI cs.CV 80%

Scaling Language-Centric Omnimodal Representation Learning

Chenghao Xiao, Hou Pong Chan, Hao Zhang, Weiwen Xu, Mahani Aljunied, Yu Rong

机构 * DAMO Academy, Alibaba Group(达摩院,阿里巴巴集团)

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);cross-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17114 2025-09-08 cs.CL cs.CV cs.LG cs.MM 80%

RAVEN: Query-Guided Representation Alignment for Question Answering over Audio, Video, Embedded Sensors, and Natural Language

Subrata Biswas, Mohammad Nur Hossain Khan, Bashima Islam

专题命中 多模态训练与对齐 :multimodal(abstract);multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.CL、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.08565 2024-12-30 cs.AI cs.CL cs.CV 80%

Baichuan-Omni Technical Report

Yadong Li, Haoze Sun, Mingan Lin, Tianpeng Li, Guosheng Dong, Tao Zhang, Bowen Ding, Wei Song, Zhenglin Cheng, Yuqi Huo, Song Chen, Xu Li, Da Pan, Shusen Zhang, Xin Wu, Zheng Liang, Jun Liu, Tao Zhang, Keer Lu, Yaqi Zhao, Yanjun Shen, Fan Yang, Kaicheng Yu, Tao Lin, Jianhua Xu, Zenan Zhou, Weipeng Chen

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);omni-modal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.17871 2024-11-06 cs.CV cs.AI cs.CL 80%

Seeing the Image: Prioritizing Visual Correlation by Contrastive Alignment

Xin Xiao, Bohong Wu, Jiacong Wang, Chunyuan Li, Xun Zhou, Haoyuan Guo

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);image-text(abstract);分类 cs.CV、cs.CL、cs.AI

Comments NeurlPS 2024, Camera ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.11172 2023-05-19 cs.CV cs.CL cs.SD eess.AS 80%

ONE-PEACE: Exploring One General Representation Model Toward Unlimited Modalities

Peng Wang, Shijie Wang, Junyang Lin, Shuai Bai, Xiaohuan Zhou, Jingren Zhou, Xinggang Wang, Chang Zhou

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);image-text(abstract);分类 cs.CV、cs.CL、eess.AS

Comments 30 pages, 9 figures, 18 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.23242 2026-08-25 cs.MM 新提交 79%

Mind the Couch! Eliciting MLLM Reasoning in Interior Design via Weak-to-Strong Task Vector Injection

留意沙发!通过弱到强任务向量注入激发多模态大语言模型在室内设计中的推理能力

Yuxuan Yang, Jingyao Wang, Luntian Mou

专题命中 多模态训练与对齐 :MLLM(title);multimodal(abstract);分类 cs.MM

AI总结 针对MLLMs在室内设计中因模态对齐问题产生的空间碰撞与审美不和谐,本文提出DART-I机制,通过弱到强任务向量注入引导MLLMs推理,无需微调即可实现精确推理,且经实验验证有效。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.22979 2026-08-25 cs.AI 新提交 79%

SA-RSQ: A Versatile Sparse Representation Framework for Multi-modal Recommender Systems

SA-RSQ:一种适用于多模态推荐系统的通用稀疏表示框架

Xiang Wang, Shigang Quan, Tingzhen Chang, Kang Yang, Sitong Chen, Yabo Fan, Xingxing Wang, Zhaodian He

机构 * Tianjin University(天津大学) Meituan(美团) Institute of Software Chinese Academy of Sciences(中国科学院软件研究所)

专题命中 多模态训练与对齐 :multi-modal(title);multimodal(abstract);分类 cs.AI

AI总结 针对工业多模态推荐系统的存储与延迟开销问题,提出SA-RSQ稀疏表示框架,经实验验证其在重构性能与CTR间的权衡表现良好,在线测试实现CTR与CPM的相对提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.21881 2026-08-25 cs.CV 新提交 79%

Region-Weighted Losses and Model Fusion for Cross-Modal PET Attenuation Correction

面向跨模态PET衰减校正的区域加权损失与模型融合

Khoa Tuan Nguyen, Joris Vankerschaver, Wesley De Neve

机构 * Center for Biosystems and Biotech Data Science, Ghent University Global Campus(根特大学全球校区生物系统与生物技术数据科学中心) Ghent University(根特大学)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

AI总结 针对BIC-MAC挑战赛,采用区域加权Carney衰减系数空间L₁损失、结合未配准DIXON MRI输入,通过两个独立模型的凸组合实现跨模态PET衰减校正,在公共验证排行榜总体排名第一。

Comments ntkhoa team submission for BIC-MAC MICCAI26 challenge (https://www.codabench.org/competitions/12555/#/results-tab)

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.17949 2026-08-25 cs.CV 版本更新 79%

SkyNative: A Native Multimodal Architecture for Remote Sensing Vision-Language Understanding

SkyNative: 一种面向遥感视觉证据推理的原生多模态框架

Xiao Yang, Ronghao Fu, Zhiwen Lin, Zhuoran Duan, Lang Sun, Jiaqi Liu, Jiashun Zhu, Jiasen Hu, Xu Na, Bo Yang

机构 * College of Computer Science and Technology, Jilin University, China(吉林大学计算机科学与技术学院) Key Laboratory of Symbolic Computation and Knowledge Engineering of Ministry of Education(教育部符号计算与知识工程重点实验室)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出SkyNative,一种原生多模态框架,通过去除预训练视觉骨干,直接在语言模型token空间中表示图像为原始patch tokens,以提升遥感图像的细粒度空间推理能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.20969 2026-08-24 cs.CV 新提交 79%

Kinematic Knowledge Maps for Pattern Alignment: Structured Latent Representational Learning in Multimodal Gait Analysis

用于模式对齐的运动学知识图谱:多模态步态分析中的结构化潜在表示学习

Chen Dong, He Zonglin, Cheung Kenneth M. C

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本研究提出ScoliDetect框架,通过运动学知识图谱(KKM)实现多模态步态分析,其介导的融合结合三模态对比预训练提升了脊柱侧凸筛查的泛化性与可解释性,外部ROC-AUC达0.972。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.19669 2026-08-21 cs.CV cs.LG 新提交 79%

Scaffolding Minds: Optimizing Latent Visual Target Representations for Multimodal Reasoning

搭建心智:优化多模态推理的潜在视觉目标表示

Haoqiang Kang, Yinpeng Chen, Luyang Liu, Jesper Sparre Andersen, Abhijit Ogale, Baochen Sun, Lichan Hong, Ed H. Chi

机构 * Google DeepMind(谷歌DeepMind) UC San Diego(加州大学圣迭戈分校)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 针对潜在推理框架的 SFT 阶段潜在表示对齐差、RL 阶段缺乏探索性潜在轨迹的局限,提出 Scaffolding Minds,学习专用搭建编码器与 RL 采样器的均值方差,在多基准任务上取得显著性能提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07687 2026-08-21 eess.IV cs.CV 版本更新 79%

FermatSyn: SAM2-Enhanced Bidirectional Mamba with Isotropic Spiral Scanning for Multi-Modal Medical Image Synthesis

FermatSyn: 基于改进双向Mamba的多模态医学图像合成方法

Feng Yuan, Yifan Gao, Haoyue Li, Xin Gao

机构 * USTC(中国科学技术大学) SII(上海信息研究所)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

AI总结 FermatSyn通过改进的双向Mamba结合Fermat螺旋扫描策略,解决多模态医学图像合成中全局一致性与局部细节的平衡问题,提升合成图像质量与临床应用价值。

Comments MICCAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.19063 2026-08-20 cs.CV 新提交 79%

When Two Tracers Disagree: An Investigation of Multimodal Fusion for Clinical PET/CT Segmentation

当两种示踪剂意见不合时:临床PET/CT分割的多模态融合研究

Jack A. Johnson, Bartłomiej W. Papież

机构 * University of Oxford(牛津大学) Nuffield Department of Medicine(纳菲尔德医学系) Department of Oncology(肿瘤学系) Big Data Institute(大数据研究所) Nuffield Department of Population Health(纳菲尔德人口健康系)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 该研究针对前列腺癌PSMA与FDG PET/CT的多模态融合分割,训练示踪剂特异性3D nnU-Net基线并对比多种融合策略,发现融合未始终优于单示踪剂模型,需更优架构保留示踪剂特异性表征

Comments 10 pages (8 pages main content and 2 pages of references), 2 figures, 2 tables, accepted to MICCAI 2026 Cancer Prevention, Detection, and IntervenTion (CaPTion) Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏