InstructX: Towards Unified Visual Editing with MLLM Guidance
机构 * Intelligent Creation Team, ByteDance(字节跳动智能创作团队)
专题命中 多模态生成 :MLLM(title,abstract);multimodal(abstract);分类 cs.CV
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
机构 * Intelligent Creation Team, ByteDance(字节跳动智能创作团队)
专题命中 多模态生成 :MLLM(title,abstract);multimodal(abstract);分类 cs.CV
机构 * City University of Hong Kong(香港城市大学) ; WeChat, Tencent Inc(微信、腾讯公司) ; Manycore Tech Inc(很多核科技公司)
专题命中 多模态生成 :multimodal(title);multi-modal(abstract);分类 cs.CV
Comments Code, model, and dataset will be released at project page soon: https://luckyhzt.github.io/x2video
机构 * Show Lab, National University of Singapore(展示实验室,新加坡国立大学)
专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI
Comments Project Page: https://showlab.github.io/Paper2Video/
机构 * Shanghai Jiaotong University(上海交通大学) ; Huawei Noah’s Ark Lab(华为诺亚实验室) ; East China Normal University(华东师范大学) ; Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(深圳先进技术研究院,中国科学院)
专题命中 多模态生成 :multimodal(abstract);分类 cs.CV
Comments 25pages,20figures
机构 * School of Computer Science and Engineering, Sun Yat-sen University(中山大学计算机科学与工程学院) ; Guangdong Key Laboratory of Big Data Analysis and Processing(广东大数据分析与处理重点实验室)
专题命中 多模态生成 :multi-modal(abstract);分类 cs.AI
Comments Website with code: https://guowei-zou.github.io/dm1/
机构 * Microsoft(微软公司) ; Massachusetts General Hospital, Harvard University(哈佛大学麻省总医院) ; Emory University(埃默里大学)
专题命中 多模态生成 :multi-modal(abstract);分类 cs.AI
专题命中 多模态生成 :multi-modal(abstract)