TAViS: Text-bridged Audio-Visual Segmentation with Foundation Models
TAViS: 基于文本的音频视觉分割与基础模型
机构 * Northwestern Polytechnical University(西北工业大学) ; Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) ; Institute of Artificial Intelligence, Hefei Comprehensive National Science Center(合肥综合国家科学中心人工智能研究院)
专题命中 音频语音多模态 :audio-visual(title,abstract);multimodal(abstract);cross-modal(abstract);multimodal foundation model(abstract)
AI总结 TAViS通过文本桥接机制结合多模态基础模型与分割模型,实现高效的音频视觉分割与跨模态对齐。
Comments ICCV2025,code:https://github.com/Sssssuperior/TAViS