arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-15 至 2025-10-15 共收录 10 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 10 篇

2510.12254 2025-10-15 cs.LG 82%

FedMMKT:Co-Enhancing a Server Text-to-Image Model and Client Task Models in Multi-Modal Federated Learning

Ningxin He, Yang Liu, Wei Sun, Xiaozhou Ye, Ye Ouyang, Tiegang Gao, Zehui Zhang

机构 * Institute for AI Industry Research, Tsinghua University(人工智能产业研究院,清华大学) School of Software Engineering, Nankai University(软件工程学院,南开大学) Department of Computing, Hong Kong Polytechnic University(计算机学院,香港理工大学) AsiaInfo Technologies(亚信息科技) China-Austria Belt and Road Joint Laboratory on Artificial Intelligence and Advanced Manufacturing, Hangzhou Dianzi University(人工智能与先进制造联合实验室,杭州电子科技大学)

专题命中 多模态生成 :multi-modal(title,abstract);multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01483 2025-10-15 cs.GR cs.CV 79%

GarmageNet: A Multimodal Generative Framework for Sewing Pattern Design and Generic Garment Modeling

Siran Li, Ruiyang Liu, Chen Liu, Zhendong Wang, Gaofeng He, Yong-Lu Li, Xiaogang Jin, Huamin Wang

机构 * Zhejiang Sci-Tech University Style3D Research Hangzhou China State Key Lab of CAD\&CG, Zhejiang University Shanghai Jiao Tong University Shanghai China State Key Lab of CAD\&CG, Zhejiang University Hangzhou China Style3D Research Shanghai Jiao Tong University

专题命中 多模态生成 :multimodal(title);multi-modal(abstract);分类 cs.CV

Comments 23 pages,20 figures

Journal ref ACM Trans. Graph. 44, 6, Article 216 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12789 2025-10-15 cs.CV cs.AI cs.LG 73%

UniFusion: Vision-Language Model as Unified Encoder in Image Generation

Kevin Li, Manuel Brack, Sudeep Katakol, Hareesh Ravi, Ajinkya Kale

机构 * Adobe Applied Research(Adobe应用研究)

专题命中 多模态生成 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments Project page at https://thekevinli.github.io/unifusion/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12325 2025-10-15 cs.IR cs.AI 70%

Causal Inspired Multi Modal Recommendation

Jie Yang, Chenyang Gu, Zixuan Liu

机构 * National University of Singapore Master of Industrial and Systems Engineering(新加坡国立大学工业与系统工程硕士) East China Normal University Master of Library and Information Science(华东师范大学图书馆与信息科学硕士) Shandong Normal University(山东师范大学)

专题命中 多模态生成 :multimodal(abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12000 2025-10-15 cs.SD cs.CL cs.LG 70%

UALM: Unified Audio Language Model for Understanding, Generation and Reasoning

Jinchuan Tian, Sang-gil Lee, Zhifeng Kong, Sreyan Ghosh, Arushi Goel, Chao-Han Huck Yang, Wenliang Dai, Zihan Liu, Hanrong Ye, Shinji Watanabe, Mohammad Shoeybi, Bryan Catanzaro, Rafael Valle, Wei Ping

机构 * CMU(卡内基梅隆大学) NVIDIA(英伟达) UMD(马里兰大学)

专题命中 多模态生成 :multimodal(abstract);cross-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12777 2025-10-15 cs.CV 57%

What If : Understanding Motion Through Sparse Interactions

Stefan Andreas Baumann, Nick Stracke, Timy Phan, Björn Ommer

机构 * Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

Comments Project page and code: https://compvis.github.io/flow-poke-transformer

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.10742 2025-10-15 cs.AI 57%

The Philosophical Foundations of Growing AI Like A Child

Dezhi Luo, Yijiang Li, Hokin Deng

机构 * University of Michigan(密歇根大学) University of California San Diego(加州大学圣地亚哥分校) Carnegie Mellon University(卡内基梅隆大学)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12253 2025-10-15 cs.LG cs.AI 57%

Diffusion Models for Reinforcement Learning: Foundations, Taxonomy, and Development

Changfu Xu, Jianxiong Guo, Yuzhu Liang, Haiyang Huang, Haodong Zou, Xi Zheng, Shui Yu, Xiaowen Chu, Jiannong Cao, Tian Wang

机构 * Jiangxi University of Finance and Economics(江西财经大学) Beijing Normal University(北京师范大学) Anhui University(安徽大学) Macquarie University(麦考瑞大学) University of Technology Sydney(悉尼大学) The University of Science and Technology (Guangzhou)(广州大学) The Hong Kong Polytechnic University(香港理工大学)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.AI

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12095 2025-10-15 cs.CV 57%

IL3D: A Large-Scale Indoor Layout Dataset for LLM-Driven 3D Scene Generation

Wenxu Zhou, Kaixuan Nie, Hang Du, Dong Yin, Wei Huang, Siqiang Guo, Xiaobo Zhang, Pengbo Hu

机构 * University of Science and Technology of China(中国科学技术大学) Songying Technology(宋英科技)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments 9 pages main paper; 15 pages references and appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07880 2025-10-15 cs.NI eess.SP 50%

Generative Resource Allocation for 6G O-RAN with Diffusion Policies

Salar Nouri, Mojdeh Karbalaeimotaleb, Vahid Shah-Mansouri, Tarik Taleb

专题命中 多模态生成 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏