LaDiC: Are Diffusion Models Really Inferior to Autoregressive Counterparts for Image-to-Text Generation?
专题命中 多模态生成 :image-text(abstract);分类 cs.CV、cs.CL、cs.AI
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
专题命中 多模态生成 :image-text(abstract);分类 cs.CV、cs.CL、cs.AI
专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI、cs.MM
Comments Accepted to CVPR 2024. Project website: http://facebookresearch.github.io/IllustratedInstructions. Code reproduction: https://github.com/sachit-menon/generating-illustrated-instructions-reproduction
专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI、cs.MM
专题命中 多模态生成 :any-to-any(abstract);分类 cs.CV、cs.CL、cs.AI
Comments Code and models are available at https://github.com/dvlab-research/MiniGemini
专题命中 多模态生成 :image-text(abstract);分类 cs.CV、cs.CL、cs.AI
Comments First two authors contributed equally; Project website: https://selma-t2i.github.io/
专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI
Comments 10 pages of main paper, 4 pages of appendix; 10 figures in main paper, 3 figures in appendix
专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL、cs.MM
Comments Accepted at ACM TOMM; ACM Transactions on Multimedia Computing, Communications, and Applications
专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI
Comments ACL 2023 (Main)
Journal ref Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pages 9174-9193
专题命中 多模态生成 :cross-modal(abstract);分类 cs.CL、cs.AI、eess.AS
Comments Demo and implementation at https://auffusion.github.io
专题命中 多模态生成 :cross-modal(abstract);分类 cs.CL、cs.AI、eess.AS
Comments ACL 2023. arXiv admin note: substantial text overlap with arXiv:2305.07760
专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.AI、cs.MM
Journal ref Published at NeurIPS 2023
专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI
专题命中 多模态生成 :image-text(abstract);分类 cs.CV、cs.CL、cs.AI
Comments Accepted to NeurIPS 2023
专题命中 多模态生成 :cross-modal(abstract);分类 cs.AI、cs.MM、eess.AS
Comments Accepted by ICML 2023. Demo and implementation at https://audioldm.github.io. Evaluation toolbox at https://github.com/haoheliu/audioldm_eval
专题命中 多模态生成 :image-text(abstract);分类 cs.CV、cs.AI、cs.MM
专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.AI、cs.MM
Comments Accepted by ACM MM'23
专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV、cs.CL、cs.AI
Comments Accepted to ICCV 2023, Project Page: https://promptstyler.github.io/
专题命中 多模态生成 :cross-modal(abstract);分类 cs.AI、cs.MM、eess.AS
Comments 16 pages, 3 figures, 2 tables, demo page: https://musicldm.github.io/
专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI
Comments Equal contribution: Bingshuai Liu and Longyue Wang. Work done while Bingshuai Liu and Chengyang Lyu were interning at Tencent AI Lab. Zhaopeng Tu is the corresponding author
专题命中 多模态生成 :cross-modal(abstract);分类 cs.CL、cs.AI、cs.MM
Comments Presented at AAAI23 CreativeAI workshop (Non-Archival). A later version is accepted to ACL23
专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI
Comments Accepted to Robotics: Science and Systems (RSS) 2023. The previous version appeared in CoRL Workshop on Language and Robot Learning 2022
专题命中 多模态生成 :image-text(abstract);分类 cs.CV、cs.AI、cs.MM
Comments This paper has 29 pages with 22 figures, including rich supplementary information. Project page is at \url{https://classifier-as-generator.github.io/}
专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI
Comments This paper is current under peer-review in IEEE TNNLS
专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI
专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI
Comments To appear at Findings of ACL 2022
专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.MM
Comments ACM MM 2021 (Video and Demo Track). Code: https://github.com/researchmm/generate-it
专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL、cs.MM
Comments The first two authors contributed to this work equally
专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV、cs.AI、cs.MM
Comments Accepted to TMM, an extended version of a paper published in ACM MM 2019. arXiv admin note: substantial text overlap with arXiv:1908.00999
专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI
Comments ICML 2021 (15 pages, 4 figures, 14 tables)
专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI
Comments Accepted for MICCAI 2019