发表机构
XJTU-POLIMI Joint School, Xi’an Jiaotong University; University of California, Berkeley; Shenzhen Future Network of Intelligence Institute (FNii-Shenzhen)(西安交通大学-米兰理工大学联合学校(西安交通大学); 加州大学伯克利分校; 深圳未来智能网络研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究诊断出流式情感理解中存在先前信念污染问题,提出无需训练的EmoUpdate框架缓解该问题,在多语音语言模型和基准测试中均实现准确率的显著提升。
AI 中文摘要
流式情感理解在持续解读当前音频时会利用历史状态,常将模型的先前预测反馈作为上下文。我们发现这种历史条件作用会扭曲当前感知:在平衡的CREMA-D-Stream反事实诊断中,仅改变注入的先前情感标签而保持音频固定,当前音频的准确率从72.50%降至30.42%,65.69%的预测发生翻转;该效应具有强标签不对称性,先验拉动幅度在4.76%至98.20%之间,我们将此缺陷命名为先前信念污染(PBC)。为解决PBC,我们提出EmoUpdate,这是一种无需训练的框架,通过三个组件分离当前音频感知与历史状态修正:(1)先验盲声学防火墙,阻止历史状态进入感知;(2)证据缩减因果信念过滤器,仅在观测形成后引入历史,且仅当有观测证据支持时保留标签不对称转移结构;(3)闭式去污染算子,源自相同的反事实测量,适用于无法使用防火墙的场景。在四个SpeechLMs和两个流式情感基准测试中,EmoUpdate在所有8种模型-基准设置下均取得最佳步长准确率和状态平衡准确率,较最强的受控基线,S-BAcc提升最高达69.71个百分点,步长准确率提升最高达38.41个百分点。
英文摘要
Streaming emotion understanding uses historical state while continuously interpreting current audio, often feeding the model's previous prediction back as context. We show that this history conditioning can distort current perception. On a balanced CREMA-D-Stream counterfactual diagnostic, changing only the injected previous emotion label while holding the audio fixed reduces current-audio accuracy from 72.50% to 30.42% and flips 65.69% of predictions. The effect is strongly label-asymmetric, with prior pull ranging from 4.76% to 98.20%, revealing a failure we call previous-belief contamination (PBC). To address PBC, we introduce EmoUpdate, a training-free framework that separates current-audio perception from historical state revision through three components: (1) a prior-blind acoustic firewall that prevents historical state from entering perception; (2) an evidence-shrunk causal belief filter that introduces history only after observation formation and retains label-asymmetric transition structure only when supported by observed evidence; and (3) a closed-form decontamination operator derived from the same counterfactual measurements for serving stacks where firewalling is unavailable. Across four SpeechLMs and two streaming emotion benchmarks, EmoUpdate achieves the best step accuracy and state-balanced accuracy in all eight model--benchmark settings, improving S-BAcc by up to 69.71 points and step accuracy by up to 38.41 points over the strongest controlled baselines.