arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-08-27 至 2025-08-27 共收录 6 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 6 篇

2508.11433 2025-08-27 cs.CV 83%

MM-R1: Unleashing the Power of Unified Multimodal Large Language Models for Personalized Image Generation

Qian Liang, Yujia Wu, Kuncheng Li, Jiwei Wei, Shiyuan He, Jinyu Guo, Ning Xie

专题命中 多模态生成 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.15194 2025-08-27 cs.CV cs.AI cs.LG 81%

DiffBlender: Composable and Versatile Multimodal Text-to-Image Diffusion Models

Sungnyun Kim, Junsoo Lee, Kibeom Hong, Daesik Kim, Namhyuk Ahn

机构 * KAIST AI(韩国科学技术院人工智能研究所) NAVER WEBTOON AI Sookmyung Women’s University(成均馆女子大学) Inha University(釜山大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Expert Systems with Applications 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18421 2025-08-27 cs.CV 70%

Why Relational Graphs Will Save the Next Generation of Vision Foundation Models?

Fatemeh Ziaeetabar

机构 * Department of Computer Science, School of Mathematics, Statistics and Computer Science, College of Science, University of Tehran, Tehran, Iran(塔里斯坦大学科学学院计算机科学系)

专题命中 多模态生成 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15761 2025-08-27 cs.CV 57%

Waver: Wave Your Way to Lifelike Video Generation

Yifu Zhang, Hao Yang, Yuqi Zhang, Yifei Hu, Fengda Zhu, Chuang Lin, Xiaofeng Mei, Yi Jiang, Bingyue Peng, Zehuan Yuan

机构 * Bytedance Waver Team(字节跳动Waver团队)

专题命中 多模态生成 :MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06905 2025-08-27 cs.CV 57%

MultiRef: Controllable Image Generation with Multiple Visual References

Ruoxi Chen, Dongping Chen, Siyuan Wu, Sinan Wang, Shiyun Lang, Petr Sushko, Gaoyang Jiang, Yao Wan, Ranjay Krishna

机构 * Zhejiang Wanli University(浙江万里大学) University of Washington(华盛顿大学) Huazhong University of Science and Technology(华中科技大学) Allen Institute for AI(人工智能研究院)

专题命中 多模态生成 :image-text(abstract);分类 cs.CV

Comments Accepted to ACM MM 2025 Datasets

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11639 2025-08-27 cs.LG 50%

Deep Generative Methods and Tire Architecture Design

Fouad Oubari, Raphael Meunier, Rodrigue Décatoire, Mathilde Mougeot

机构 * ENS Paris-Saclay, Centre Borelli(巴黎-萨克雷大学ENS分校,Borelli中心) Université Paris-Saclay, CNRS, ENS Paris-Saclay, Centre Borelli(巴黎-萨克雷大学,法国国家科学研究中心,巴黎-萨克雷大学ENS分校,Borelli中心) ENSIIE(Évry法国的ENSIEE)

专题命中 多模态生成 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏