发表机构
William & Mary(威廉与玛丽学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对基于LLM的文本到语音模型,引入两种情感敏感度量并借助激活修补干预,识别出稀疏的源到读出组件级情感控制电路,该电路结合共享与情感特定组件,能显著恢复或抑制情感读出偏移并影响解码语音的声学特征,揭示了从参考信息到生成语音情感属性的紧凑因果路径。
AI 中文摘要
基于LLM的文本到语音(TTS)模型能够生成富有情感的语音,但参考情感如何通过模型路由并在解码语音中实现仍不清楚。我们引入了两个针对匹配的中性和情感合成的度量——码本轨迹得分和晚期残差方向得分——并用它们来评分激活修补干预。在受控的匹配参考条件下,该分析识别出一个稀疏的源到读出组件级电路:每种情感涉及23至27个注意力头和MLP,约占所考虑组件的5%,在保留案例上恢复或抑制了晚期情感读出偏移的74%至88%。该电路结合了共享组件主干和情感特定组件;跨情感激活交换在48例中的47例中减少了目标读出。在解码语音中,相同的干预在每种情感的24个匹配对上产生了一致的音高、能量和频谱亮度变化。一个读出匹配的残差方向基线仅产生干预音高效果的17%至27%,表明仅内部读出移动并不能解释解码的声学变化。这些结果追踪了一条从参考派生的前缀信息到生成语音的情感相关属性的紧凑因果路径。
英文摘要
LLM-based text-to-speech (TTS) models can generate emotionally expressive speech, but how reference emotion is routed through the model and realized in decoded speech remains unclear. We introduce two emotion-sensitive metrics for matched neutral and emotional syntheses---a codec trajectory score and a late residual direction score---and use them to score activation-patching interventions. Under controlled matched-reference conditions, this analysis identifies a sparse source-to-readout component-level circuit: 23--27 attention heads and MLPs per emotion, roughly 5% of the components considered, recover or suppress 74--88% of the late emotion-readout shift on held-out cases. The circuit combines a shared component backbone with emotion-specific components; cross-emotion activation swaps reduce the target readout in 47 of 48 cases. In decoded speech, the same intervention produces consistent changes in pitch, energy, and spectral brightness over 24 matched pairs per emotion. A readout-matched residual-direction baseline produces only 17--27% of the intervention's pitch effect, showing that internal readout movement alone does not explain the decoded acoustic changes. These results trace a compact causal route from reference-derived prefix information to emotion-relevant properties of generated speech.
CommentsAccepted to EMNLP 2026 (Main Conference). 15 pages, 4 figures