AI 中文总结
本文提出掩码时间情感锚定(M-TAG)训练目标,将音频情感识别重构为时间情感锚定任务,以解决标签时间模糊问题,并验证了现有模型在此任务上的局限性。
AI 中文摘要
音频情感识别(AER)通常为整个录音分配单一标签,当存在多个说话人和情感事件时,该标签的时间范围变得模糊。我们通过将AER重新表述为时间情感锚定(TAG)任务来解决这一局限性,该任务将情感与时间上有界的语音片段和声调描述相关联。为支持这一表述,我们整理了现有情感识别数据集的时间标注版本,并构建了包含两到四个情感语音片段(包括重叠语音)的录音。在这种较长格式下使用标准语言建模目标进行训练会降低情感识别和时间锚定性能,而声调描述可能为情感预测提供捷径。为应对这些挑战,我们引入了掩码时间情感锚定(M-TAG),这是一种监督训练目标,结合了全序列语言建模与注意力掩码下的情感和时间戳交叉熵损失。掩码改变了情感标签标记可见的上下文,以减少对捷径的依赖并提高泛化能力,而时间戳损失则结合了距离感知权重以惩罚较大的时间误差。我们评估了EMO-TAG(使用我们的数据集和目标微调的模型)在情感识别和情感时间锚定方面的性能,并与三个AER和音频语言基线(Flamingo-Next、Audio-Reasoner和AffectGPT)进行比较。我们的结果表明,现有模型尽管情感识别性能具有竞争力,但在情感时间锚定方面能力有限。
英文摘要
Audio emotion recognition (AER) typically assigns a single label to an entire recording, leaving the temporal scope of that label ambiguous when multiple speakers and affective events are present. We address this limitation by reformulating AER as a Temporal Affective Grounding (TAG) task that associates emotions with temporally bounded speech spans and vocal tone descriptions. To support this formulation, we curate temporally annotated versions of existing emotion recognition datasets and construct recordings containing two to four affective speech spans, including overlapping speech. Training in this longer format with a standard language-modeling objective can degrade both emotion recognition and temporal grounding performance, while tone descriptions can provide shortcuts for emotion prediction. To address these challenges, we introduce Masked Temporal Affective Grounding (M-TAG), a supervised training objective that combines full-sequence language modeling with emotion and timestamp cross-entropy losses under attention masking. The masking varies the context visible to emotion-label tokens to reduce reliance on shortcuts and improve generalization, while the timestamp loss incorporates a distance-aware weight to penalize larger temporal errors. We evaluate EMO-TAG, a model fine-tuned using our dataset and objective, on emotion recognition and affective temporal-grounding against three AER and audio-language baselines: Flamingo-Next, Audio-Reasoner, and AffectGPT. Our results show that existing models achieve limited affective temporal-grounding despite competitive emotion recognition performance.