arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4965 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4965 篇

2512.20436 2025-12-24 eess.IV cs.AI cs.CV 81%

Dual-Encoder Transformer-Based Multimodal Learning for Ischemic Stroke Lesion Segmentation Using Diffusion MRI

基于双编码变压器的多模态学习用于扩散磁共振成像的缺血性中风病变分割

Muhammad Usman, Azka Rehman, Muhammad Mutti Ur Rehman, Abd Ur Rehman, Muhammad Umar Farooq

机构 * Department of Anesthesiology, Perioperative and Pain Medicine, Stanford University(麻醉学、围术期与疼痛医学系,斯坦福大学) Department of Biomedical Sciences, Seoul National University(生物医学科学系,首尔国立大学) Department of Computer Engineering, National University of Sciences and Technology (NUST)(计算机工程系,国立科学与技术大学(NUST)) Department of Computer Science, The University of Alabama(计算机科学系,阿拉巴马大学) Department of Computer Science, Hanyang University(计算机科学系,翰阳大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 本文提出双编码变压器架构,利用多模态扩散MRI实现缺血性中风病变分割,达到85.4%的Dice分数。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.01298 2025-12-23 cs.CV cs.AI 81%

Towards Enhanced Image Generation Via Multi-modal Chain of Thought in Unified Generative Models

通过统一生成模型中的多模态推理链提升图像生成

Yi Wang, Mushui Liu, Wanggui He, Hanyang Yuan, Longxiang Zhang, Ziwei Huang, Guanghao Zhang, Wenkai Fang, Haoze Jiang, Shengxuming Zhang, Dong She, Jinlong Liu, Weilong Dai, Mingli Song, Hao Jiang, Jie Song

机构 * Zhejiang University(浙江大学) Alibaba Group(阿里巴巴集团)

专题命中 多模态生成 :multi-modal(title);multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本文通过引入多模态推理链提升统一生成模型的复杂图像生成能力,提出FoXperts架构和MCoT方法,实现更高效的多模态生成。

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.14862 2025-12-22 cs.LG cs.CL cs.CV 81%

LatentExplainer: Explaining Latent Representations in Deep Generative Models with Multimodal Large Language Models

LatentExplainer: 通过多模态大语言模型解释深度生成模型中的潜在表示

Mengdan Zhu, Raasikh Kanjiani, Jiahui Lu, Andrew Choi, Qirui Ye, Liang Zhao

机构 * Emory University(埃默里大学) University College London(伦敦大学学院)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL

AI总结 LatentExplainer通过多模态大语言模型为深度生成模型的潜在变量生成语义解释,提升模型可解释性。

Comments Accepted to CIKM 2025 Full Research Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15069 2025-12-18 cs.CV cs.AI 81%

PMMD: A pose-guided multi-view multi-modal diffusion for person generation

PMMD: 一种基于姿态的多视角多模态扩散用于人物生成

Ziyu Shang, Haoran Liu, Rongchao Zhang, Zhiqian Wei, Tongtong Feng

机构 * Harbin Institute of Technology, Shenzhen, China(哈尔滨工业大学(深圳)) City University of Hong Kong, Hong Kong, China(香港城市大学) Peking University, Beijing, China(北京大学) Tsinghua University, Beijing, China(清华大学)

专题命中 多模态生成 :multi-modal(title);multimodal(abstract);分类 cs.CV、cs.AI

AI总结 PMMD通过多视角多模态扩散框架,实现可控姿态和外观的人像生成,提升一致性、细节保留和可控性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13752 2025-12-17 cs.CV cs.AI 81%

STAR: STacked AutoRegressive Scheme for Unified Multimodal Learning

STAR:用于统一多模态学习的堆叠自回归方案

Jie Qin, Jiancheng Huang, Limeng Qiao, Lin Ma

机构 * Meituan Inc(美团公司)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 STAR通过堆叠自回归方案提升多模态生成性能,同时保持理解能力,实验验证其在统一多模态学习中的有效性。

Comments 18 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11464 2025-12-15 cs.CV cs.AI cs.LG 81%

Exploring MLLM-Diffusion Information Transfer with MetaCanvas

探索MLLM-扩散信息传输与MetaCanvas

Han Lin, Xichen Pan, Ziqi Huang, Ji Hou, Jialiang Wang, Weifeng Chen, Zecheng He, Felix Juefei-Xu, Junzhe Sun, Zhipeng Fan, Ali Thabet, Mohit Bansal, Chu Wang

机构 * Meta Superintelligence Labs(Meta 超智能实验室) New York University(纽约大学) Nanyang Technological University(南洋理工大学)

专题命中 多模态生成 :MLLM(title);multimodal(abstract);分类 cs.CV、cs.AI

AI总结 MetaCanvas通过使MLLMs在潜在空间中推理和规划,提升多模态生成的精确度和结构化控制。

Comments Project page: https://metacanvas.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09610 2025-12-11 cs.HC cs.AI cs.CV 81%

ImageTalk: Designing a Multimodal AAC Text Generation System Driven by Image Recognition and Natural Language Generation

ImageTalk: 设计一种由图像识别和自然语言生成驱动的多模态AAC文本生成系统

Boyin Yang, Puming Jiang, Per Ola Kristensson

机构 * Department of Engineering, University of Cambridge(剑桥大学工程系) Department of Computing, Imperial College London(伦敦帝国学院计算机系)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 ImageTalk通过图像识别和自然语言生成设计多模态AAC文本生成系统,实现95.6%的按键节省率和高用户满意度。

Comments 24 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.09854 2025-12-04 cs.CL cs.AI cs.LG 81%

Scaling Multimodal Search and Recommendation with Small Language Models via Upside-Down Reinforcement Learning

通过倒置强化学习扩展小语言模型以支持多模态搜索与推荐

Yu-Chen Lin, Sanat Sharma, Hari Manikandan, Jayant Kumar, Tracy Holloway King, Jing Zheng

机构 * Adobe(Adobe公司) Meta

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL、cs.AI

AI总结 本文通过倒置强化学习和合成数据蒸馏,利用小语言模型实现高效多模态搜索与推荐,显著降低推理延迟和内存开销。

Comments Accepted by ICDM 2025 MMSR

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02351 2025-12-03 cs.CV cs.AI 81%

Understanding and Harnessing Sparsity in Unified Multimodal Models

理解并利用统一多模态模型中的稀疏性

Shwai He, Chaorui Deng, Ang Li, Shen Yan

机构 * ByteDance Seed(字节跳动种子) University of Maryland, College Park(马里兰大学学院公园分校)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 本研究通过MoE适应方法,利用稀疏激活提升统一多模态模型的效率,使模型在激活约一半参数的情况下达到与完整模型相当的性能。

Comments 13 pages, 13 figures, 8 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02088 2025-12-03 eess.IV cs.AI cs.CV cs.LG 81%

Comparing Baseline and Day-1 Diffusion MRI Using Multimodal Deep Embeddings for Stroke Outcome Prediction

比较基线和第1天扩散磁共振成像用于中风预后预测的多模态深度嵌入

Sina Raeisadigh, Myles Joshua Toledo Tan, Henning Müller, Abderrahmane Hedjoudje

机构 * 1 Department of Computer Science, University of Geneva, Switzerland 2 Department of Electrical \& Computer Engineering, University of Florida, FL, USA 3 Service of Medical Informatics, University Hospital of Geneva, Switzerland 4 Department of Imaging Medical Informatics, University of Geneva, Switzerland

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 本研究通过多模态深度嵌入方法,利用基线和治疗后1天的扩散MRI数据,结合临床特征和病变体积,预测急性缺血性中风患者3个月的功能预后。

Comments 5 pages, 5 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12170 2025-12-02 cs.CV cs.AI 81%

Rethinking Multimodal Point Cloud Completion: A Completion-by-Correction Perspective

重新思考多模态点云补全:一种补全-修正视角

Wang Luo, Di Wu, Hengyuan Na, Yinlin Zhu, Miao Hu, Guocong Quan

机构 * Wang Luo, Di Wu, Hengyuan Na, Yinlin Zhu, Miao Hu, Guocong Quan(作者)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 PGNet通过补全-修正范式,结合多阶段框架和双特征编码,实现更鲁棒的点云补全,提升重建精度与结构一致性。

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00677 2025-12-02 cs.CV cs.AI 81%

Dynamic-eDiTor: Training-Free Text-Driven 4D Scene Editing with Multimodal Diffusion Transformer

Dynamic-eDiTor: 基于文本的无训练4D场景编辑与多模态扩散变换器

Dong In Lee, Hyungjun Doh, Seunggeun Chi, Runlin Duan, Sangpil Kim, Karthik Ramani

机构 * Purdue University(普渡大学) Korea University(韩国大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 Dynamic-eDiTor通过多模态扩散变换器和4DGS实现无训练文本驱动的4D场景编辑,提升多视图和时间一致性。

Comments 4D Scene Editing

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.20561 2025-12-02 cs.CV cs.CL 81%

Does Understanding Inform Generation in Unified Multimodal Models? From Analysis to Path Forward

在统一多模态模型中,理解是否会影响生成?从分析到未来路径

Yuwei Niu, Weiyang Jin, Jiaqi Liao, Chaoran Feng, Peng Jin, Bin Lin, Zongjian Li, Bin Zhu, Weihao Yu, Li Yuan

机构 * Peking University(北京大学) Chongqing University(重庆大学) HKU MMLab(香港大学多模态实验室) PengCheng Laboratory(鹏城实验室)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL

AI总结 研究通过UniSandbox分析统一多模态模型中理解与生成之间的差距,发现显式链式思维和自训练方法能有效弥合这一差距,并揭示查询架构的潜在CoT特性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10125 2025-12-02 cs.CV cs.MM 81%

Proxy-Tuning: Tailoring Multimodal Autoregressive Models for Subject-Driven Image Generation

代理调优:为以主题驱动的图像生成定制多模态自回归模型

Yi Wu, Shengju Qian, Lingting Zhu, Lei Liu, Wandi Qiao, Ziqiang Li, Lequan Yu, Bin Li

机构 * University of Science and Technology of China(中国科学技术大学) The Chinese University of Hong Kong(香港中文大学) The University of Hong Kong(香港大学) Nanjing University of Information Science and Technology(南京信息工程大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.MM

AI总结 本文提出代理调优方法,通过扩散模型增强AR模型在主题驱动图像生成中的能力,揭示了弱到强泛化现象,提升了多主题组合和上下文理解的表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.18428 2025-11-26 cs.AI cs.CL 81%

Multi-Modal Data Exploration via Language Agents

通过语言代理进行多模态数据探索

Farhad Nooralahzadeh, Yi Zhang, Jonathan Furst, Kurt Stockinger

机构 * Zurich University of Applied Sciences(瑞士应用科学大学)

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CL、cs.AI

AI总结 本文提出M$^2$EX系统,通过语言代理实现多模态数据探索,优于现有系统在准确性和性能指标上

Comments Accepted to the IJCNLP AACL 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.14993 2025-11-26 cs.AI cs.CV 81%

Multi-modal Generative AI: Multi-modal LLMs, Diffusions, and the Unification

多模态生成式AI:多模态大语言模型、扩散模型与统一

Xin Wang, Yuwei Zhou, Bin Huang, Hong Chen, Wenwu Zhu

机构 * Department of Computer Science, Beijing Information Science and Technology National Research Center, Tsinghua University(计算机系,北京信息科学与技术国家研究中心,清华大学)

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV、cs.AI

AI总结 本文综述了多模态生成式AI,涵盖多模态LLMs、扩散模型及统一模型的设计与应用,探讨了统一理解和生成的方法及未来研究方向。

Comments 21 pages, 10 figures, 3 tables

Journal ref IEEE Transactions on Circuits and Systems for Video Technology 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03227 2025-11-07 cs.HC cs.AI cs.MM 81%

Node-Based Editing for Multimodal Generation of Text, Audio, Image, and Video

Alexander Htet Kyaw, Lenin Ravindranath Sivalingam

机构 * Massachusetts Institute of Technology(麻省理工学院) Microsoft Research(微软研究院)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.AI、cs.MM

Comments Accepted to NeurIPS 2025, Conference on Neural Information Processing Systems, Workshop on Generative and Protective AI for Content Creation

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21448 2025-11-06 eess.AS cs.CV cs.SD 81%

ThinkSound: Chain-of-Thought Reasoning in Multimodal Large Language Models for Audio Generation and Editing

Huadai Liu, Kaicheng Luo, Jialei Wang, Wen Wang, Qian Chen, Zhou Zhao, Wei Xue

机构 * Hong Kong University of Science and Technology (HKUST)(香港理工大学) Tongyi Fun Team, Alibaba Group(阿里云团队) Zhejiang University(浙江大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、eess.AS

Comments Accepted by NeurIPS 2025 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19755 2025-11-04 cs.LG cs.AI cs.CV 81%

A Survey on Cache Methods in Diffusion Models: Toward Efficient Multi-Modal Generation

Jiacheng Liu, Xinyu Wang, Yuqi Lin, Zhikai Wang, Peiru Wang, Peiliang Cai, Qinming Zhou, Zhengan Yan, Zexuan Yan, Zhengyi Shi, Chang Zou, Yue Ma, Linfeng Zhang

机构 * Shanghai Jiao Tong University(上海交通大学) Tsinghua University(清华大学) The Hong Kong University of Science and Technology(香港科技大学)

专题命中 多模态生成 :multi-modal(title);multimodal(abstract);分类 cs.CV、cs.AI

Comments 22 pages,2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27632 2025-11-03 cs.CV cs.AI 81%

Sketch-to-Layout: Sketch-Guided Multimodal Layout Generation

Riccardo Brioschi, Aleksandr Alekseev, Emanuele Nevali, Berkay Döner, Omar El Malki, Blagoj Mitrevski, Leandro Kieliger, Mark Collier, Andrii Maksai, Jesse Berent, Claudiu Musat, Efi Kokiopoulou

机构 * EPFL(苏黎世联邦理工学院) Google DeepMind(谷歌DeepMind)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 15 pages, 18 figures, GitHub link: https://github.com/google-deepmind/sketch_to_layout, accept at ICCV 2025 Workshop (HiGen)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26105 2025-10-31 cs.CV cs.AI cs.CR 81%

Security Risk of Misalignment between Text and Image in Multi-modal Model

Xiaosen Wang, Zhijin Ge, Shaokang Wang

机构 * Xidian University(西安电子科技大学) Shanghai Jiaotong University(上海交通大学)

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24514 2025-10-29 cs.CV cs.CL 81%

Latent Sketchpad: Sketching Visual Thoughts to Elicit Multimodal Reasoning in MLLMs

Huanyu Zhang, Wenshan Wu, Chengzu Li, Ning Shang, Yan Xia, Yangyu Huang, Yifan Zhang, Li Dong, Zhang Zhang, Liang Wang, Tieniu Tan, Furu Wei

机构 * MSR(微软研究院) UCAS(中国科学院自动化研究所) CASIA(中国科学院自动化研究所) Cambridge(剑桥大学) NJU(南京大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22684 2025-10-28 cs.CV cs.CL 81%

RoboSVG: A Unified Framework for Interactive SVG Generation with Multi-modal Guidance

Jiuniu Wang, Gongjie Zhang, Quanhao Qian, Junlong Gao, Deli Zhao, Ran Xu

机构 * DAMO Academy, Alibaba Group(达摩院,阿里巴巴集团)

专题命中 多模态生成 :multi-modal(title);multimodal(abstract);分类 cs.CV、cs.CL

Comments 15 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22521 2025-10-28 cs.CV cs.AI cs.IR cs.LG 81%

Open Multimodal Retrieval-Augmented Factual Image Generation

Yang Tian, Fan Liu, Jingyuan Zhang, Wei Bi, Yupeng Hu, Liqiang Nie

机构 * Shandong University(山东大学) National University of Singapore(新加坡国立大学) Kuaishou Technology(快手科技) Harbin Institute of Technology(哈尔滨工业大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.18400 2025-10-22 eess.IV cs.AI cs.CV 81%

A Multimodal Deep Learning Approach for White Matter Shape Prediction in Diffusion MRI Tractography

Yui Lo, Yuqian Chen, Dongnan Liu, Leo Zekelman, Jarrett Rushmore, Yogesh Rathi, Nikos Makris, Alexandra J. Golby, Fan Zhang, Weidong Cai, Lauren J. O'Donnell

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Paper accepted to Human Brain Mapping. 25 pages, 3 figures, 8 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09121 2025-10-21 cs.CV cs.AI 81%

MSDM: Generating Task-Specific Pathology Images with a Multimodal Conditioned Diffusion Model for Cell and Nuclei Segmentation

Dominik Winter, Mai Bui, Monica Azqueta Gavaldon, Nicolas Triltsch, Marco Rosati, Nicolas Brieu

机构 * AstraZeneca Computational Pathology GmbH(阿斯利康计算病理学 GmbH)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13253 2025-10-20 cs.CV cs.AI cs.LG 81%

End-to-End Multi-Modal Diffusion Mamba

Chunhao Lu, Qiang Lu, Meichen Dong, Jake Luo

机构 * China University of Petroleum-Beijing(中国石油大学(北京)) Leyard Optoelectronic(莱亚德光电) University of Wisconsin-Milwaukee(威斯康星大学密尔沃基分校)

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.08209 2025-10-17 cs.CV cs.AI cs.LG 81%

Emergent Visual Grounding in Large Multimodal Models Without Grounding Supervision

Shengcao Cao, Liang-Yan Gui, Yu-Xiong Wang

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments ICCV 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09479 2025-10-14 cs.AI cs.CL 81%

Draw with Thought: Unleashing Multimodal Reasoning for Scientific Diagram Generation

Zhiqing Cui, Jiahao Yuan, Hanqing Wang, Yanshu Li, Chenxu Du, Zhenglong Ding

机构 * Nanjing University of Information Science \& Technology Nanjing China East China Normal University Shanghai China The Hong Kong University of Science Brown University Providence America Southwest Jiaotong University Chengdu China Nanjing University of Information Science \& Technology East China Normal University Brown University Southwest Jiaotong University

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments 10 pages, 5 figures, accepted to appear in the Proceedings of the 33rd ACM International Conference on Multimedia (MM '25)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21787 2025-10-13 cs.CV cs.CL 81%

DeHate: A Stable Diffusion-based Multimodal Approach to Mitigate Hate Speech in Images

Dwip Dalal, Gautam Vashishtha, Anku Rani, Aishwarya Reganti, Parth Patwa, Mohd Sarique, Chandan Gupta, Keshav Nath, Viswanatha Reddy, Vinija Jain, Aman Chadha, Amitava Das, Amit Sheth, Asif Ekbal

机构 * MIT Media Lab, USA(麻省理工学院媒体实验室) Stanford University, USA(斯坦福大学) Amazon GenAI, USA(亚马逊生成人工智能) University of South Carolina, USA(南卡罗来纳大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Defactify 3 workshop at AAAI 2024

详情

展开后加载摘要…

URL PDF HTML 收藏