大规模语言模型中的用法调制情感表示
Usage-Modulated Sentiment Representations in Large Language Models
浏览论文内容
中文总结 AI 辅助
本研究证明大型语言模型中情感表示不仅包含共享极性方向,还受用法因素调制,残差结构可被擦除和引导,为情感控制提供新方法。
中文摘要 AI 辅助
先前的研究表明,情感通常可以通过大型语言模型激活空间中的近似线性方向来捕捉,但单一方向可能无法完全捕捉情感表示。在自然交流中,情感不仅由极性塑造,还受用法因素(如语气和受众适应)的影响。我们测试这些因素是否在共享情感方向之外系统性地调制情感表示。我们构建了一个受控配对数据集,在保持事件内容固定的同时变化情感极性和用法因素,并分析了Llama、Mistral和Gemma模型。我们识别出一个共享的情感方向,将其移除,并通过擦除和生成时语气引导测试残差结构。在各模型中,共享方向是稳健的(中位余弦相似度为0.953-0.975),但移除它后,原始正负表示差异范数的0.833-0.909仍然保留。残差包含紧凑、可复现的用法条件结构。针对性的擦除比随机和标签打乱的对照更削弱了保留的用法指标。在Llama上,沿残差化语气分量引导的输出在92.8%的盲目标语气比较中被偏好,同时在98.7%的评估输出中保持了所请求的情感极性。
英文摘要
Prior work suggests that sentiment can often be captured by approximately linear directions in LLM activation spaces, but a single direction may not fully capture sentiment representations. In natural communication, sentiment is shaped not only by polarity but also by usage factors, such as tone and audience adaptation. We test whether these factors systematically modulate sentiment representations beyond a shared sentiment direction. We construct a controlled paired dataset that holds event content fixed while varying sentiment polarity and usage factors, and analyze Llama, Mistral, and Gemma. We identify a shared sentiment direction, remove it, and test the residual structure through erasure and generation-time tone steering. Across models, the shared direction is robust (median cosine 0.953-0.975), yet removing it leaves 0.833-0.909 of the original positive-negative representation-difference norm. The residuals contain compact, reproducible usage-conditioned structure. Targeted erasure weakens held-out usage metrics more than random and label-shuffled controls. On Llama, outputs steered along residualized tone components are preferred in 92.8% of blind target-tone comparisons while preserving the requested sentiment polarity in 98.7% of evaluated outputs.
发表机构
- William & Mary(威廉与玛丽学院)
机构由 AI 辅助整理,请以论文原文为准。