通过结构化说话人条件实现基于Whisper的稳健以婴儿为中心的多层音频理解
Robust Multi-Tier Infant-Centered Audio Understanding with Whisper via Structured Speaker Conditioning
浏览论文内容
中文总结 AI 辅助
针对以婴儿为中心的自然音频标注难题,本文提出结合LoRA微调Whisper与轻量Transformer的多层音频标注器,通过序列级平滑损失与分解式说话人token设计提升稳健性,可高效完成家庭全天音频的婴儿相关标注。
中文摘要 AI 辅助
近期模型设计与自监督音频表征的进展提升了语音与音频理解能力,但以婴儿为中心的自然录音因标注数据有限、信噪比低及跨家庭域偏移而仍具挑战性。本文提出一种家庭条件下的多层音频标注器,它结合经LoRA微调的Whisper编码器与轻量型目标说话人感知Transformer,用于跨层级的长上下文推理与逐帧预测。为提升时间一致性,我们引入简单的序列级平滑损失;为增强跨家庭的稳健性,我们提出带共享层级token与可学习家庭特定偏移的分解式说话人token设计,以此减少家庭偏差并促进可泛化表征。这些设计共同实现了对家庭环境下全天音频记录的高效且有效的以婴儿为中心的音频标注。
英文摘要
Recent advances in model design and self-supervised audio representations have improved speech and audio understanding, yet infant-centered naturalistic recordings remain challenging due to limited labeled data, low signal-to-noise ratio, and cross-family domain shifts. We present a family-conditioned, multi-tier audio tagger that combines a LoRA-finetuned Whisper encoder with a lightweight, target-speaker-aware Transformer for long-context inference and framewise prediction across tiers. To improve temporal coherence, we incorporate a simple sequence-level smoothing loss, and to enhance robustness across households, we introduce a factorized speaker-token design with a shared tier token and a learned family-specific offset, reducing family bias and promoting generalizable representations. Together, these choices enable efficient and effective infant-centered audio tagging of daylong audio recordings in home environments.
发表机构
- University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
- University of Arizona(亚利桑那大学)
- Worcester Polytechnic Institute(伍斯特理工学院)
机构由 AI 辅助整理,请以论文原文为准。