Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation
机构 * Character AI ; Yale University(耶鲁大学)
专题命中 多模态生成 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.MM、eess.AS
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
机构 * Character AI ; Yale University(耶鲁大学)
专题命中 多模态生成 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.MM、eess.AS
专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL
Comments This paper is accepted for presentation in TRB annual meeting 2026. The version presented here is the preprint version before peer review process
机构 * National University of Singapore(新加坡国立大学) ; Shanghai Jiao Tong University(上海交通大学) ; Shanghai AI Lab(上海人工智能实验室)
专题命中 多模态生成 :multi-modal(title)
机构 * School of Vehicle and Mobility & College of AI, Tsinghua University(车辆与移动学院及人工智能学院,清华大学) ; School of Artificial Intelligence, University of Chinese Academy of Sciences(人工智能学院,中国科学院大学) ; School of Mechanical Engineering, University of Science and Technology Beijing(机械工程学院,北京科技大学)
专题命中 多模态生成 :multimodal(abstract);分类 cs.AI
专题命中 多模态生成 :multi-modal(abstract);分类 cs.AI
Comments Fixed and extended results
专题命中 多模态生成 :multimodal(abstract);分类 cs.CV
Comments NeurIPS 2024. Project page at https://diffcut-segmentation.github.io. Code at https://github.com/PaulCouairon/DiffCut
专题命中 多模态生成 :multimodal(abstract)
Comments 14 pages, to be published at the 26th International Conference on Artificial Intelligence in Education (AIED '25)