VA-DPO:用于语言模型可控情感生成的效价-唤醒直接偏好优化
VA-DPO: Valence-Arousal Direct Preference Optimization for Controllable Emotion Generation in Language Models
浏览论文内容
中文总结 AI 辅助
该研究针对现有情感生成无法表达细粒度情感的问题,提出VA-DPO方法,通过效价-唤醒连续目标优化LoRA适配器,在多模型上实现细粒度情感生成且不损害通用性能。
中文摘要 AI 辅助
我们能多精确地告诉语言模型该产生何种情感?大多数情感生成研究采用离散标签(如快乐、愤怒、悲伤),无法表达“轻度低落但平静”这类目标情感。我们转而将期望情感指定为效价-唤醒平面中的连续点(v*, a*),并训练模型使其达到该点。我们的方法VA-DPO是对直接偏好优化(DPO)的小幅修改:一个冻结的VA回归器根据采样生成内容与目标的欧氏距离打分,仅保留距离差超过阈值τ的候选对,再通过普通DPO损失优化LoRA适配器(对比冻结参考模型)。DPO目标本身未变,新的是偏好数据的构建方式。在Llama-3.1-8B-Instruct上,该方法使平均VA距离较系统提示降低33%,较少样本提示降低25%,效价/唤醒相关性提升至r_v=0.93和r_a=0.75。该增益可迁移至Qwen3-8B和Llama-3.2-3B,且无常规代价:MMLU性能不变(Δ=+0.0),HellaSwag和TruthfulQA性能得以保留。我们发布了代码、配置及偏好构建流程。
英文摘要
How precisely can we tell a language model how to feel? Most work on emotional generation answers with a discrete label - happy, angry, sad - which cannot express a target like "mildly downcast but calm." We instead specify the desired affect as a continuous point (v*, a*) in the Valence-Arousal plane and train the model to hit it. Our method, VA-DPO, is a small modification to Direct Preference Optimization: a frozen VA regressor scores each sampled generation by its Euclidean distance to the target, we keep only candidate pairs whose distance gap clears a margin tau, and we optimize a LoRA adapter with the ordinary DPO loss against a frozen reference. The DPO objective itself is unchanged; what is new is how the preference data is built. On Llama-3.1-8B-Instruct this cuts mean VA distance to the target by 33% over system-prompting and 25% over few-shot prompting, lifting valence/arousal correlation to r_v=0.93 and r_a=0.75. The gains carry over to Qwen3-8B and Llama-3.2-3B, and they do not come at the usual price: MMLU is unchanged (Delta=+0.0) and HellaSwag and TruthfulQA are preserved. We release the code, configs, and the preference-construction pipeline.