A Multimodal Framework for Understanding Collaborative Design Processes
专题命中 音频语音多模态 :multimodal(title,abstract)
Comments Accepted to IEEE VIS 2025
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
专题命中 音频语音多模态 :multimodal(title,abstract)
Comments Accepted to IEEE VIS 2025
机构 * Department of MetaBioHealth, Sungkyunkwan University, Korea(韩国成均馆大学代谢生物健康系) ; Department of Applied Artificial Intelligence, Sungkyunkwan University, Korea(韩国成均馆大学应用人工智能系) ; Department of Computer Science and Engineering, Jaume I University, Spain(西班牙伊萨贝拉大学计算机科学与工程系)
专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.AI、cs.MM
Comments Accepted for publication at IJCAI 2025. 9 pages, 4 tables, 3 figures
机构 * Nanjing University of Science
专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV
Comments Accepted for publication at ACMMM 2025
机构 * Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) ; Huawei Technologies Co., Ltd.(华为技术有限公司)
专题命中 音频语音多模态 :any-to-any(abstract);分类 cs.AI
Comments Accepted by INTERSPEECH 2025