DiffThinker: Towards Generative Multimodal Reasoning with Diffusion Models
DiffThinker: 向基于扩散模型的生成多模态推理迈进
机构 * Shanghai AI Laboratory(上海人工智能实验室) ; Nanjing University(南京大学) ; Shanghai Jiao Tong University(上海交通大学) ; The Chinese University of Hong Kong(香港中文大学)
专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV
AI总结 DiffThinker通过基于扩散模型的生成方法,在多模态推理任务中实现了更高效的视觉推理和更精确的空间处理,显著优于现有模型。
Comments Project page: https://diffthinker-project.github.io