SpeechCT-CLIP: Distilling Text-Image Knowledge to Speech for Voice-Native Multimodal CT Analysis
机构 * Computer Aided Medical Procedures, Technical University of Munich(慕尼黑技术大学计算机辅助医学程序)
专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL、eess.AS
Comments Submitted to ICASSP 2026; under review