arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-09-12 至 2025-09-12 共收录 4 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4 篇

2509.09456 2025-09-12 cs.CV 83%

FlexiD-Fuse: Flexible number of inputs multi-modal medical image fusion based on diffusion model

Yushen Xu, Xiaosong Li, Yuchun Wang, Xiaoqi Cheng, Huafeng Li, Haishu Tan

机构 * School of Physics(物理学院) Optoelectronic Engineering, Foshan University(光电工程学院,佛山大学) Guangdong-HongKong-Macao Joint Laboratory for Intelligent Micro-Nano Optoelectronic Technology(粤港澳联合智能微纳光电技术实验室) Guangdong Provincial Key Laboratory of Industrial Intelligent Inspection Technology(广东省工业智能检测技术重点实验室) School of Information Engineering(信息工程学院) Automation, Kunming University of Science(自动化系,昆明理工大学)

专题命中 多模态生成 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

Journal ref Expert Systems with Applications, 2025: 128895

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08847 2025-09-12 cs.AI cs.CL cs.LG cs.SE 81%

Automated Unity Game Template Generation from GDDs via NLP and Multi-Modal LLMs

Amna Hassan

机构 * UET Taxila(塔希尔大学工程学院)

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20877 2025-09-12 cs.CV 74%

Deep Learning Framework for Early Detection of Pancreatic Cancer Using Multi-Modal Medical Imaging Analysis

Dennis Slobodzian, Amir Kordijazi

专题命中 多模态生成 :multi-modal(title);分类 cs.CV

Comments 21 pages, 17 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08775 2025-09-12 cs.RO 50%

Joint Model-based Model-free Diffusion for Planning with Constraints

Wonsuhk Jung, Utkarsh A. Mishra, Nadun Ranawaka Arachchige, Yongxin Chen, Danfei Xu, Shreyas Kousik

机构 * Georgia Institute of Technology(佐治亚理工学院)

专题命中 多模态生成 :multi-modal(abstract)

Comments The first two authors contributed equally. Last three authors advised equally. Accepted to CoRL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏