发表机构
Independent Researcher(独立研究者)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究在开源全双工语音模型Moshi中,通过均值差异激活转向探索情感控制,发现情感可线性解码但转向效果不均,快乐、愤怒、惊讶共享方向,悲伤独特可转向。
AI 中文摘要
全双工语音代理需要在实时对话中调节情感和表达方式,例如在缓和投诉时、在调度中传达紧迫感、在缓和临床结果时。情感和表达控制已在TTS和基于回合的模型中通过提示条件合成、参考条件合成和激活转向得到了充分研究;PersonaPlex在双工模型中控制了身份,但未涉及情感。我们研究了Moshi(一个完全开源的全双工语音语言模型)中四种情感的情感转向,使用均值差异激活转向,该方法每帧仅需几次向量加法,无需重新训练。我们表明,情感可以从Moshi的残差流中线性解码,但激活转向仅部分可实现,且不均匀;快乐、愤怒和惊讶转向共享方向,而悲伤则明显可转向。我们还表明,这三种情感共享的成分不能简单地从所有情感中同等投影去除。
英文摘要
Full-duplex voice agents need to modulate emotion and delivery during real-time conversations, when de-escalating a complaint, carrying urgency in dispatch, softening a clinical result. Emotion and delivery control is well studied for TTS and turn based models through prompt-conditioned synthesis, reference-conditioned synthesis and activation steering; PersonaPlex controls identity in a duplex model but not affect. We study emotion steering in Moshi, a fully open sourced full-duplex speech language model, across four emotions, using mean-difference activation steering, which costs only a few vector additions per frame and no retraining. We show that emotion is linearly decodable from Moshi's residual stream, but activation steering is only partially achievable, and unevenly so; as happy, angry and surprise steer towards a shared direction while sad is distinctly steerable. We also show that the shared component across the three emotions cannot simply be projected away from all the emotions equally.
CommentsAccepted at the NeurIPS 2026 Workshop on Real-Time Conversational Agents (RTCA), Sydney. OpenReview: https://openreview.net/forum?id=6IUNRorKmj