你无法偏好未采样的情绪:DPO微调大语言模型中的强度欠冲
You Can't Prefer Emotions You Don't Sample: Intensity Undershoot in DPO-Tuned LLMs
浏览论文内容
中文总结 AI 辅助
研究发现DPO微调LLM在情感强度上系统欠冲,归因于候选池缺乏极端样本,通过均匀采样更热更大的候选池可提升效价增益并降低外推误差。
中文摘要 AI 辅助
让语言模型以“非常兴奋”的语气回应,其输出通常只是稍微更有活力。我们量化了这一效应。我们将指令微调的大语言模型(LLM)条件化在一个连续的效价-唤醒(VA)目标上,其中效价衡量状态令人愉悦的程度,唤醒衡量其激活程度,并使用一个冻结的回归器测量所达到的情感,将请求的目标从-1扫到+1。响应移动的幅度远小于要求:增益,即达到的情感对请求的情感的斜率,在Llama-3.1-8B上对效价仅为0.26,对唤醒仅为0.13,而一个忠实的控制器应得分为1。模型系统性地欠冲请求的情感强度,这为Fazzi等人(2025)的定性观察提供了量化数据。我们的实验将这一现象追溯到偏好学习流程。来自自然语料库(如EmoBank)的训练目标以中性为主,且采样的候选本身很少达到极端情感,因此直接偏好优化(DPO)没有极端示例可供偏好。如果我们均匀覆盖目标空间并采样一个更“热”、更大的候选池,效价增益从0.26升至0.40±0.02(3个随机种子),外推误差下降,仅以适度的分布内代价(EmoBank测试VA距离从0.092升至0.107)为代价。同样的方法在Qwen3-8B上复现(效价增益0.44,分布内精度保持)。唤醒更难且更不可靠:其增益平均几乎不动,并在种子间波动(0.14±0.07,而效价波动为±0.02),因为提高唤醒需要基础模型不愿生成的候选。证据表明,忠实的强度受限于候选池的极端性,而非条件化格式。
英文摘要
Ask a language model to respond "very excitedly," and its output is typically only mildly more energetic. We quantify this effect. We condition an instruction-tuned LLM on a continuous Valence-Arousal (VA) target, where valence measures how pleasant a state is and arousal how activated it is, measure the achieved affect with a frozen regressor, and sweep the requested target from -1 to +1. The response moves far less than asked: the gain, the slope of achieved against requested affect, is only 0.26 for valence and 0.13 for arousal on Llama-3.1-8B, where a faithful controller would score 1. The model systematically undershoots requested emotional intensity, which puts a number on the qualitative observation of Fazzi et al. (2025). Our experiments trace this to the preference-learning pipeline. Training targets from natural corpora such as EmoBank are neutral-heavy, and the sampled candidates themselves rarely reach extreme affect, so Direct Preference Optimization (DPO) is left with no extreme exemplar to prefer. If instead we cover the target space uniformly and sample a hotter, larger candidate pool, valence gain rises from 0.26 to 0.40 +/- 0.02 (3 seeds) and extrapolation error drops, at only a modest in-distribution cost (EmoBank-test VA distance 0.092 to 0.107). The same recipe reproduces on Qwen3-8B (gain_v 0.44, with in-distribution accuracy preserved). Arousal is harder and less reliable: its gain barely moves on average and swings across seeds (0.14 +/- 0.07, against valence's tight +/- 0.02), because raising arousal needs candidates the base model is reluctant to generate. The evidence indicates that faithful intensity is bottlenecked by the extremity of the candidate pool rather than by the conditioning format.
发表机构
- Hanyang University(汉阳大学)
机构由 AI 辅助整理,请以论文原文为准。