基于大语言模型增强的音频-文本对齐的零样本呼吸声分类
Zero-Shot Respiratory Sound Classification through LLM-Augmented Audio-Text Alignment
- Eindhoven University of Technology(埃因霍温理工大学)
- Singapore Management University(新加坡管理大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究提出LLM增强的音频-文本对齐框架,解决呼吸编码器零样本推理的临床语义缺失问题,在9项任务上的零样本AUC优于CLAP等模型,数据效率更高。
AI中文摘要:
自监督呼吸编码器缺乏零样本推理所需的临床领域语义基础,在无特定任务标注数据时实用性受限。本文提出一种框架,将这些编码器与医学术语在共享潜在空间中对齐,使其转变为具备零样本能力的基础模型。为应对配对数据稀缺问题,我们使用医学大语言模型从元数据中合成结构化报告,为对比学习创建密集语义锚点。训练过程结合基于sigmoid的对比损失、编码器原生自监督学习目标及感知相似性的负采样,以强化病理边界。在6个数据集的9项任务中,本方法实现61.3%的平均零样本AUC,超越CLAP(51.4%)和Qwen2-Audio(54.9%),且仅使用全规模基线模型43%的数据便达到最高线性探测AUC(71.6%),表明结构化语义对齐在临床诊断中优于大规模通用模型。
英文摘要:
Self-supervised respiratory encoders lack semantic grounding in clinical domain needed for zero-shot inference, limiting their utility without task-specific labeled data. We propose a framework that aligns these encoders with medical terminology in a shared latent space turning them into a zero-shot-capable foundation model. To address paired data scarcity, we use a medical LLM to synthesize structured reports from metadata, creating dense semantic anchors for contrastive learning. Our training combines a sigmoid-based contrastive loss with encoder's native SSL objective and similarity-aware negative sampling to sharpen pathological boundaries. Across 9 tasks on 6 datasets, our method achieves a 61.3% mean zero-shot AUC, surpassing CLAP (51.4%) and Qwen2-Audio (54.9%) while reaching the highest linear probing AUC (71.6%) with only 43% of data used by full-scale baselines, showing that structured semantic alignment outperforms large-scale, general-purpose models in clinical diagnostics.