arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-11-17 至 2025-11-17 共收录 4 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4 篇

2511.11066 2025-11-17 cs.CV cs.AI cs.CL 83%

S2D-ALIGN: Shallow-to-Deep Auxiliary Learning for Anatomically-Grounded Radiology Report Generation

Jiechao Gao, Chang Liu, Yuangang Li

专题命中 多模态生成 :multimodal(abstract);multi-modal(abstract);cross-modal(abstract);image-text(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07375 2025-11-17 cs.CV 79%

Novel Diffusion Models for Multimodal 3D Hand Trajectory Prediction

Junyi Ma, Wentao Bao, Jingyi Xu, Guanzhong Sun, Xieyuanli Chen, Hesheng Wang

机构 * IRMV Lab, the Department of Automation, Shanghai Jiao Tong University(IRMV实验室,自动化系,上海交通大学) Meta Reality Labs(Meta现实实验室) the Department of Electronic Engineering, Shanghai Jiao Tong University(电子工程系,上海交通大学) the School of Information and Control Engineering, China University of Mining and Technology(信息与控制工程学院,中国矿业大学) the College of Intelligence Science and Technology, National University of Defense Technology(智能科学与技术学院,国防科技大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments Accepted to IROS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11434 2025-11-17 cs.CV 70%

WEAVE: Unleashing and Benchmarking the In-context Interleaved Comprehension and Generation

Wei Chow, Jiachun Pan, Yongyuan Liang, Mingze Zhou, Xue Song, Liyu Jia, Saining Zhang, Siliang Tang, Juncheng Li, Fengda Zhang, Weijia Wu, Hanwang Zhang, Tat-Seng Chua

机构 * National University of Singapore(新加坡国立大学) Nanyang Technological University(南洋理工大学) University of Maryland, College Park(马里兰大学学院公园分校) Zhejiang University(浙江大学)

专题命中 多模态生成 :multimodal(abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07334 2025-11-17 cs.CV 57%

Unleashing the Potential of Large Language Models for Text-to-Image Generation through Autoregressive Representation Alignment

Xing Xie, Jiawei Liu, Ziyue Lin, Huijie Fan, Zhi Han, Yandong Tang, Liangqiong Qu

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV

Comments Accepted by AAAI 2026 Oral

详情

展开后加载摘要…

URL PDF HTML 收藏