arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-08-18 至 2025-08-18 共收录 3 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 3 篇

2508.11159 2025-08-18 cs.LG 78%

Mitigating Modality Quantity and Quality Imbalance in Multimodal Online Federated Learning

Heqiang Wang, Weihong Yang, Xiaoxiong Zhong, Jia Zhou, Fangming Liu, Weizhe Zhang

机构 * Peng Cheng Laboratory(鹏城实验室)

专题命中 多模态生成 :multimodal(title,abstract)

Comments arXiv admin note: text overlap with arXiv:2505.16138

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11153 2025-08-18 cs.CV 57%

LEARN: A Story-Driven Layout-to-Image Generation Framework for STEM Instruction

Maoquan Zhang, Bisser Raytchev, Xiujuan Sun

机构 * Graduate School of Advanced Science and Engineering, Hiroshima University(Hiroshima大学研究生院) Department of Computer Science, Weifang University of Science and Technology(潍坊科技大学计算机科学系)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments The International Conference on Neural Information Processing (ICONIP) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.05908 2025-08-18 cs.CV 57%

GBR: Generative Bundle Refinement for High-fidelity Gaussian Splatting with Enhanced Mesh Reconstruction

Jianing Zhang, Yuchao Zheng, Ziwei Li, Qionghai Dai, Xiaoyun Yuan

机构 * College of future information technology, Fudan University(未来信息科技学院,复旦大学) School of Biomedical Engineering, Tsinghua University(生物医学工程学院,清华大学) Key Laboratory for Information Science of Electromagnetic Waves (MoE), Fudan University(电磁波信息科学重点实验室(MoE),复旦大学) Department of Automation, Tsinghua University(自动化系,清华大学) MoE Key Lab of Artificial Intelligence, AI Institute, Shanghai Jiao Tong University(人工智能重点实验室,人工智能研究院,上海交通大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏