发表机构
Instituto Politécnico Nacional (IPN); Centro de Investigación en Computación (CIC)(墨西哥国立理工学院; 计算机研究中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对RLHF引发摘要情感漂移的问题,提出Policy Attribution框架溯源漂移来源,验证了跨语言漂移特性,并提出感知情感的正则化技术以降低漂移,且相关代码将公开。
AI 中文摘要
人类反馈强化学习(RLHF)可将大型语言模型(LLM)与人类偏好对齐,提升摘要的流畅性与安全性,但会引发情感漂移:生成的摘要过于中立,缺失情感细微差别。本文诊断了RL为何会成为情感中和器,提出Policy Attribution框架,该框架利用梯度与对数几率分解,将漂移溯源至奖励模型(RM)信号及KL(Kullback-Leibler)惩罚项。情感漂移体现为在偏好不确定性下,模型偏向选择“低风险” token以最大化期望奖励的策略性偏差(Stiennon等人,2020;Gao、Schulman与Hilton,2023)。在Reddit TL;DR和CNN/DailyMail数据集上,RLHF生成的摘要获得更高奖励,但情感方差降低30%-40%。对8种语言的跨语言分析显示,漂移具有语言独立性,且形态更丰富的语言受抑制程度更高(Krasitskii等人,2026)。本文提出并验证了一种感知情感的正则化技术,可在不损害摘要质量的情况下将漂移降低18%-22%,代码与工具包将公开。
英文摘要
Reinforcement learning with human feedback (RLHF) aligns LLMs with human preferences, improving summarization fluency and safety, but causes sentiment drift: overly neutral summaries stripped of emotional nuance. We diagnose why RL acts as a sentiment neutralizer and present Policy Attribution, a framework using gradient and logit decomposition to trace drift to reward model (RM) signals and KL (Kullback-Leibler) penalty. Sentiment drift reflects a strategic bias toward "low-risk" tokens maximizing expected rewards under preference uncertainty (Stiennon et al., 2020; Gao, Schulman, and Hilton, 2023). On Reddit TL;DR and CNN/DailyMail, RLHF summaries get higher rewards but show 30-40% lower sentiment variance. Cross-lingual analysis across eight languages shows language-independent drift, with morphologically richer languages more suppressed (Krasitskii et al., 2026). We propose and validate a sentiment-aware regularization technique reducing drift by 18-22% without harming summary quality. The code and toolkit will be public.