arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

作为频谱更新重组的偏好调整

Preference Tuning as Spectral Update Reorganization

Peiyan Zhang, Haibo Jin, Liying Kang, Haohan Wang

arXiv 2607.20438首次发表:更新:

发表机构

University of Illinois at Urbana-Champaign; Hong Kong Polytechnic University(伊利诺伊大学厄巴纳-香槟分校; 香港理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究通过频谱结构研究RLHF及偏好优化,将偏好诱导更新转化为可干预对象,发现更新呈现头-尾组织,头部主导端点偏移,揭示了偏好训练后是结构化更新重组,以及对齐增益和覆盖损失与更新组织方式的关系。

AI 中文摘要

基于偏好的训练后通常通过端点行为来理解,但产生这种行为的学习更新在很大程度上仍然不透明。我们通过其诱导参数更新的频谱结构来研究基于人类反馈的强化学习(RLHF)和相关偏好优化。通过分解有效的低秩适应(LoRA)更新并将其频谱分量作为插件模块重新加载,我们将偏好诱导的更新转化为可以隔离、重新组合和直接干预的对象。在模型家族、优化算法和监督机制中,这些更新始终呈现出频谱的头-尾组织。一个紧凑的头部早期出现并携带主要的端点偏移,而一个异质的残余尾部保留。这种分裂是功能性的而非仅仅是描述性的。插件干预表明头部解释了与基础模型的可见行为差异,而尾部单独作用较弱。跨运行重新组合进一步表明混合适配器遵循头部的来源,表明头部携带运行级求解器偏差。这种端点主导并不意味着学习充分性。仅头部学习并非无意义但无法恢复完整解,特别是在分布外行为上。仅尾部学习几乎没有可见增益,但没有尾部则无法恢复完整解。这些发现将偏好训练后重新定义为结构化更新重组而非整体行为校正,并表明对齐增益和覆盖损失与学习更新本身的组织方式相关。

英文摘要

Preference-based post-training is usually understood through endpoint behavior, yet the learned update that produces this behavior remains largely opaque. We study RLHF and related preference optimization through the spectral structure of their induced parameter updates. By decomposing effective LoRA updates and reloading their spectral components as plug-in modules, we turn preference-induced updates into objects that can be isolated, recomposed, and directly intervened on. Across model families, optimization algorithms, and supervision regimes, these updates consistently develop a spectral head--tail organization. A compact head emerges early and carries the dominant endpoint shift, while a heterogeneous residual tail remains. The split is functional rather than merely descriptive. Plug-in intervention shows that the head accounts for the visible behavioral departure from the base model, while the tail is weak in isolation. Cross-run recomposition further shows that mixed adapters follow the source of the head, indicating that the head carries run-level solver bias. This endpoint dominance does not imply learning sufficiency. Head-only learning is non-vacuous but fails to recover the full solution, especially on out-of-distribution behavior. Tail-only learning yields little visible gain, yet the full solution is not recovered without the tail. These findings recast preference post-training as structured update reorganization rather than a monolithic behavioral correction, and suggest that alignment gain and coverage loss are tied to how the learned update itself is organized.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑