arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-12-16 至 2025-12-16 共收录 7 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 7 篇

2512.12756 2025-12-16 cs.CV 88%

FysicsWorld: A Unified Full-Modality Benchmark for Any-to-Any Understanding, Generation, and Reasoning

FysicsWorld: 一个统一的全模态基准用于任意到任意的理解、生成和推理

Yue Jiang, Dingkang Yang, Minghao Han, Jinghang Han, Zizhi Chen, Yizhou Liu, Mingcheng Li, Peng Zhai, Lihua Zhang

机构 * College of Intelligent Robotics and Advanced Manufacturing(智能机器人与先进制造学院) Fudan University(复旦大学) Fysics Intelligence Technologies Co., Ltd.(菲茨智能技术有限公司)

专题命中 多模态生成 :any-to-any(title,abstract);omni-modal(abstract,comments);multimodal(abstract);cross-modal(abstract)

AI总结 FysicsWorld是一个统一的全模态基准,支持图像、视频、音频和文本之间的双向交互,通过16个主要任务和3,268个样本全面评估理解、生成和推理能力,揭示模型在多模态任务中的性能差异。

Comments The omni-modal benchmark report from Fysics AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04810 2025-12-16 cs.CV 79%

EMMA: Efficient Multimodal Understanding, Generation, and Editing with a Unified Architecture

EMMA: 一种高效的多模态理解、生成与编辑统一架构

Xin He, Longhui Wei, Jianbo Ouyang, Minghui Liao, Lingxi Xie, Qi Tian

机构 * Huawei Inc.(华为公司)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

AI总结 EMMA提出了一种高效的统一多模态架构,通过高效自编码器、通道级拼接、共享解耦网络和专家混合机制,在效率和性能上超越现有方法。

Comments Project Page: https://emma-umm.github.io/emma/

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11906 2025-12-16 cs.CV cs.LG 79%

MPath: Multimodal Pathology Report Generation from Whole Slide Images

MPath:从全切片图像生成多模态病历报告

Noorul Wahab, Nasir Rajpoot

机构 * TIA Centre, Department of Computer Science, University of Warwick, UK(沃里克大学计算机科学系TIA研究中心)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

AI总结 MPath通过多模态条件化方法,利用预训练的生物医学语言模型生成病理报告,展示了在病理报告生成中的有效性与可扩展性。

Comments Pages 4, Figures 1, Table 1

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12196 2025-12-16 cs.MM cs.CV cs.SD eess.AS 67%

AutoMV: An Automatic Multi-Agent System for Music Video Generation

AutoMV: 一种自动多智能体系统用于音乐视频生成

Xiaoxuan Tang, Xinping Lei, Chaoran Zhu, Shiyun Chen, Ruibin Yuan, Yizhi Li, Changjae Oh, Ge Zhang, Wenhao Huang, Emmanouil Benetos, Yang Liu, Jiaheng Liu, Yinghao Ma

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Nanjing University(南京大学) Queen Mary University of London(伦敦大学女王学院) Hong Kong University of Science and Technology(香港科技大学) University of Manchester(曼彻斯特大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.MM、eess.AS

AI总结 AutoMV通过多智能体协作生成完整音乐视频,优于现有方法并接近专业水准。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13690 2025-12-16 cs.CV cs.AI cs.GR cs.LG 62%

DiffusionBrowser: Interactive Diffusion Previews via Multi-Branch Decoders

DiffusionBrowser: 通过多分支解码器实现交互式扩散预览

Susung Hong, Chongjian Ge, Zhifei Zhang, Jui-Hsien Wang

机构 * University of Washington(华盛顿大学) Adobe Research(Adobe研究)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.AI

AI总结 DiffusionBrowser通过多分支解码器实现交互式视频生成预览,提升生成效率与可控性。

Comments Project page: https://susunghong.github.io/DiffusionBrowser

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13276 2025-12-16 cs.CV 57%

CogniEdit: Dense Gradient Flow Optimization for Fine-Grained Image Editing

CogniEdit: 密集梯度流优化用于细粒度图像编辑

Yan Li, Lin Liu, Xiaopeng Zhang, Wei Xue, Wenhan Luo, Yike Guo, Qi Tian

机构 * Hongkong University of Science and Technology(香港科学与技术大学) Huawei Company(华为公司)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

AI总结 CogniEdit通过密集梯度流优化实现细粒度图像编辑,结合多模态推理与动态令牌焦点重新定位,提升指令遵循与视觉质量的平衡

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21329 2025-12-16 cs.AI physics.soc-ph 57%

Active Inference AI Systems for Scientific Discovery

主动推断AI系统用于科学发现

Karthik Duraisamy

机构 * Michigan Institute for Computational Discovery & Engineering(密歇根计算发现与工程研究所) University of Michigan(密歇根大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

AI总结 本文提出主动推断AI系统,通过结合缓慢假设生成与快速验证推理,利用因果和多模态模型促进科学发现,强调人类判断在处理不确定性中的关键作用。

详情

展开后加载摘要…

URL PDF HTML 收藏