S2D-ALIGN: Shallow-to-Deep Auxiliary Learning for Anatomically-Grounded Radiology Report Generation
专题命中 多模态生成 :multimodal(abstract);multi-modal(abstract);cross-modal(abstract);image-text(abstract)
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
专题命中 多模态生成 :multimodal(abstract);multi-modal(abstract);cross-modal(abstract);image-text(abstract)
机构 * IRMV Lab, the Department of Automation, Shanghai Jiao Tong University(IRMV实验室,自动化系,上海交通大学) ; Meta Reality Labs(Meta现实实验室) ; the Department of Electronic Engineering, Shanghai Jiao Tong University(电子工程系,上海交通大学) ; the School of Information and Control Engineering, China University of Mining and Technology(信息与控制工程学院,中国矿业大学) ; the College of Intelligence Science and Technology, National University of Defense Technology(智能科学与技术学院,国防科技大学)
专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV
Comments Accepted to IROS 2025
机构 * National University of Singapore(新加坡国立大学) ; Nanyang Technological University(南洋理工大学) ; University of Maryland, College Park(马里兰大学学院公园分校) ; Zhejiang University(浙江大学)
专题命中 多模态生成 :multimodal(abstract);multi-modal(abstract);分类 cs.CV
专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV
Comments Accepted by AAAI 2026 Oral