X2I: Seamless Integration of Multimodal Understanding into Diffusion Transformer via Attention Distillation
机构 * OPPO AI Center(OPPO人工智能中心) ; Tsinghua University(清华大学)
Comments Accepted to ICCV 2025
期刊&会议
International Conference on Computer Vision · 会议 · Computer Vision
机构 * OPPO AI Center(OPPO人工智能中心) ; Tsinghua University(清华大学)
Comments Accepted to ICCV 2025
机构 * VCIP, CS, Nankai University(南开大学计算机科学与技术学院) ; NKIARI, Shenzhen Futian(深圳未来科技研究院) ; USTC(University of Science and Technology of China) ; CUHK MMLab(香港中文大学MMLab) ; VAST(中国科学院自动化研究所) ; Shanghai AI Lab(上海人工智能实验室)
Comments Accepted at ICCV 2025. Project page: https://github.com/HVision-NKU/TAR3D