arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-12-03 至 2025-12-03 共收录 8 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 8 篇

2512.02351 2025-12-03 cs.CV cs.AI 81%

Understanding and Harnessing Sparsity in Unified Multimodal Models

理解并利用统一多模态模型中的稀疏性

Shwai He, Chaorui Deng, Ang Li, Shen Yan

机构 * ByteDance Seed(字节跳动种子) University of Maryland, College Park(马里兰大学学院公园分校)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 本研究通过MoE适应方法,利用稀疏激活提升统一多模态模型的效率,使模型在激活约一半参数的情况下达到与完整模型相当的性能。

Comments 13 pages, 13 figures, 8 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02088 2025-12-03 eess.IV cs.AI cs.CV cs.LG 81%

Comparing Baseline and Day-1 Diffusion MRI Using Multimodal Deep Embeddings for Stroke Outcome Prediction

比较基线和第1天扩散磁共振成像用于中风预后预测的多模态深度嵌入

Sina Raeisadigh, Myles Joshua Toledo Tan, Henning Müller, Abderrahmane Hedjoudje

机构 * 1 Department of Computer Science, University of Geneva, Switzerland 2 Department of Electrical \& Computer Engineering, University of Florida, FL, USA 3 Service of Medical Informatics, University Hospital of Geneva, Switzerland 4 Department of Imaging Medical Informatics, University of Geneva, Switzerland

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 本研究通过多模态深度嵌入方法,利用基线和治疗后1天的扩散MRI数据,结合临床特征和病变体积,预测急性缺血性中风患者3个月的功能预后。

Comments 5 pages, 5 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02273 2025-12-03 cs.CV cs.AI 62%

Progressive Image Restoration via Text-Conditioned Video Generation

通过文本条件视频生成实现渐进式图像修复

Peng Kang, Xijun Wang, Yu Yuan

机构 * Department of Computer Science, University of Illinois Springfield(伊利诺伊大学斯普林菲尔德分校计算机科学系) School of Electrical and Computer Engineering, Purdue University(普渡大学电气与计算机工程学院)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.AI

AI总结 本研究通过微调CogVideo,利用文本条件视频生成实现图像渐进式修复,提升感知度量并展示强泛化能力。

Comments First two authors contributed equally to this work. IEEE ICNC Accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06996 2025-12-03 cs.CV cs.AI 62%

Visible Yet Unreadable: A Systematic Blind Spot of Vision Language Models Across Writing Systems

可见却不可读:跨书写系统下视觉语言模型的系统性盲区

Jie Zhang, Ting Xu, Gelei Deng, Runyi Hu, Han Qiu, Tianwei Zhang, Qing Guo, Ivor Tsang

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本文研究了视觉语言模型在跨书写系统下识别碎片化文本的鲁棒性,发现其在可见但不可读的刺激下表现下降,揭示了模型对组成先验依赖不足的结构性限制。

Comments arXiv admin note: This article has been withdrawn by arXiv administrators due to violation of arXiv policy regarding generative AI authorship

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02713 2025-12-03 cs.AI 57%

Training Data Attribution for Image Generation using Ontology-Aligned Knowledge Graphs

利用本体对齐的知识图谱训练数据归因于图像生成

Theodoros Aivalis, Iraklis A. Klampanos, Antonis Troumpoukis, Joemon M. Jose

机构 * National Centre for Scientific Research ``Demokritos''(国家科学研究中心「德莫克里特」) University of Glasgow(格拉斯哥大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

AI总结 本文提出利用本体对齐的知识图谱方法,通过多模态大语言模型提取图像中的结构化三元组,以追踪生成模型中训练数据的影响,从而提升透明度和可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02554 2025-12-03 cs.CV 57%

OmniPerson: Unified Identity-Preserving Pedestrian Generation

OmniPerson: 统一的身份保持行人生成

Changxiao Ma, Chao Yuan, Xincheng Shi, Yuzhuo Ma, Yongfei Zhang, Longkun Zhou, Yujia Zhang, Shangze Li, Yifan Xu

机构 * Beihang University(北航大学) Pengcheng Laboratory(鹏城实验室)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

AI总结 OmniPerson通过统一身份保持的行人生成管道,解决ReID中数据不足问题,实现高保真生成和身份一致性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02492 2025-12-03 cs.CV 57%

YingVideo-MV: Music-Driven Multi-Stage Video Generation

YingVideo-MV: 基于音乐的多阶段视频生成

Jiahui Chen, Weida Wang, Runhua Shi, Huan Yang, Chaofan Ding, Zihao Chen

机构 * AI Lab, GiantNetwork(人工智能实验室)

专题命中 多模态生成 :audio-visual(abstract);分类 cs.CV

AI总结 YingVideo-MV通过整合音频语义分析、时间感知扩散模型和摄像机运动控制模块,实现了基于音乐的高质量视频生成。

Comments 18 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.08096 2025-12-03 cs.CV 57%

MegaSR: Mining Customized Semantics and Expressive Guidance for Real-World Image Super-Resolution

MegaSR: 为真实世界图像超分辨率挖掘定制语义和表达引导

Xinrui Li, Jinrong Zhang, Jianlong Wu, Chong Chen, Liqiang Nie, Zhouchen Lin

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

AI总结 MegaSR通过定制语义和表达引导提升真实世界图像超分辨率的重建质量与结构一致性。

详情

展开后加载摘要…

URL PDF HTML 收藏