Towards fairer public transit: Real-time tensor-based multimodal fare evasion and fraud detection
专题命中 音频语音多模态 :multimodal(title,abstract);cross-modal(abstract)
Comments 10 pages
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
专题命中 音频语音多模态 :multimodal(title,abstract);cross-modal(abstract)
Comments 10 pages
专题命中 音频语音多模态 :multimodal(title,abstract);cross-modal(abstract)
机构 * Center for Robotics, Mines Paris - PSL University Paris, France(机器人中心,巴黎 Mines Paris - PSL 大学)
专题命中 音频语音多模态 :cross-modal(abstract);audio-visual(abstract);分类 cs.CV、eess.AS
Comments ICLR 2025 (Poster). Camera ready version. Project Page: https://amandinebtto.github.io/NeRAF; 24 pages, 13 figures
Journal ref The Thirteenth International Conference on Learning Representations, 2025
机构 * Georgia Institute of Technology(佐治亚理工学院) ; Carnegie Mellon University(卡内基梅隆大学)
专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV
Comments ICCV 2025. Project page: https://clink-chop-thud.github.io/
专题命中 音频语音多模态 :audio-visual(abstract);分类 eess.AS
机构 * Georgia Institute of Technology(佐治亚理工学院) ; Bytedance Inc.(字节跳动公司) ; Squirrel AI, USA(squirrel AI 美国分公司) ; The University of Virginia(弗吉尼亚大学) ; Cornell University(康奈尔大学)
专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV
Comments Github Repo: https://github.com/AdityaLab/MM4TSA Updated to include papers accepted by IJCAI25, KDD25, ICML25, NeurIPS25 4 figures or tables, 19 pages, 251 references
机构 * MSc Artificial Intelligence Master Thesis(人工智能硕士论文)
专题命中 音频语音多模态 :multi-modal(abstract)