OpenOmni: Advancing Open-Source Omnimodal Large Language Models with Progressive Multimodal Alignment and Real-Time Self-Aware Emotional Speech Synthesis
机构 * Shenzhen Key Laboratory for High Performance Data Mining(深圳高性能数据挖掘重点实验室) ; Shenzhen Institute of Advanced Technology(深圳先进技术研究院) ; Chinese Academy of Sciences(中国科学院) ; University of Chinese Academy of Sciences(中国科学院大学) ; Tongyi Laboratory(通义实验室) ; University of New South Wales(新南威尔士大学) ; National University of Singapore(新加坡国立大学) ; University of Science and Technology of China(中国科学技术大学) ; MoE Key Laboratory of Brain-inspired Intelligent Perception and Cognition(脑启发智能感知与认知重点实验室)
专题命中 音频语音多模态 :multimodal(title);分类 cs.CV、cs.CL