AI 中文总结
研究针对设备端语音情感识别中大型模型成本高的问题,提出自适应多教师关系蒸馏方法。通过单类支持向量机分配教师权重,用关系蒸馏损失捕捉结构,在多数据集和学生架构上优于单教师蒸馏,两组件互补增益。
AI 中文摘要
设备端语音情感识别(SER)对实时应用至关重要,但擅长SER的大型自监督模型对边缘设备来说成本过高。多教师知识蒸馏可将其压缩为轻量级学生模型,但存在两个挑战:教师可靠性因批次而异,且对数级蒸馏忽略样本间关系结构。我们提出自适应多教师关系蒸馏(AMRD)来解决这两个问题。在每个教师的对数相似性矩阵上使用单类支持向量机分配有利于更一致教师的逐批权重。关系蒸馏损失使教师和学生的相似性矩阵对齐,捕捉对数匹配遗漏的结构。在IEMOCAP和CREMA-D数据集上,跨越四种学生架构,AMRD在大多数设置下优于单教师蒸馏基线,消融实验证实两个组件都产生了互补增益。
英文摘要
On-device speech emotion recognition (SER) is critical for real-time applications, yet large self-supervised models that excel at SER are too costly for edge devices. Multi-teacher knowledge distillation can compress them into a lightweight student, but two challenges remain: teacher reliability varies across batches, and logit-level distillation ignores inter-sample relational structure. We propose Adaptive Multi-teacher Relational Distillation (AMRD) to address both. A one-class SVM on each teacher's logit similarity matrix assigns per-batch weights favoring more coherent teachers. A relational distillation loss aligns teacher and student similarity matrices, capturing structure that logit matching misses. On IEMOCAP and CREMA-D datasets across four student architectures, AMRD outperforms single-teacher distillation baselines in most settings, and ablations confirm both components yield complementary gains.