arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

仅使用上下文是不够的:面向个性化奖励建模的测试时训练

Using Context Is Not Enough: Test-Time Training for Personalized Reward Modeling

Bohao Wang, Xiaoyan Zhao, Yang Zhang, Jinghang Guo, Chun Chen, Can Wang, Jiawei Chen

arXiv 2609.35109首次发表:更新:

发表机构

Zhejiang University; National University of Singapore(浙江大学; 新加坡国立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对上下文学习在个性化奖励建模中无法捕捉偏好关系的问题,提出偏好对齐的测试时训练方法,通过序列级更新快速权重,高效提升个性化奖励预测性能。

AI 中文摘要

基于人类反馈的强化学习(RLHF)使大型语言模型(LLMs)与人类偏好保持一致,然而大多数流程学习的是一个单一的奖励模型,忽略了偏好的个体差异。个性化奖励模型(PRMs)通过将奖励条件化为用户特定的反馈来解决这一问题,最常见的方式是通过上下文学习(ICL),即把用户的历史比较作为上下文偏好对提供给模型。然而,我们发现了基于ICL的PRM的一个关键局限性:它们无法捕捉上下文对所传达的偏好关系。为了解决这一问题,我们提出了偏好对齐的测试时训练(P-TTT),该方法将这些关系显式编码到用户特定的快速权重中,以进行个性化奖励预测。P-TTT引入了序列级更新和应用操作,以匹配偏好反馈的响应级粒度,同时采用偏好对齐的目标函数,直接利用成对偏好关系来指导快速权重的适应。值得注意的是,P-TTT实现简单且计算高效,在单次前向传播中更新快速权重,无需推理时的反向传播。大量实验表明,P-TTT能更有效地捕捉历史偏好关系,并以较大优势超越了最先进的方法。

英文摘要

Reinforcement learning from human feedback (RLHF) aligns large language models (LLMs) with human preferences, yet most pipelines learn a single reward model that overlooks individual differences in preferences. Personalized reward models (PRMs) address this by conditioning rewards on user-specific feedback, most commonly through in-context learning (ICL), where a user's historical comparisons are supplied as contextual preference pairs. However, we identify a key limitation of ICL-based PRMs: they fail to capture the preference relations conveyed by contextual pairs. To address this, we propose Preference-Aligned Test-Time Training (P-TTT), which explicitly encodes these relations into user-specific fast weights for personalized reward prediction. P-TTT introduces sequence-level update and apply operations to match the response-level granularity of preference feedback, together with a preference-aligned objective that directly uses pairwise preference relations to guide fast-weight adaptation. Notably, P-TTT is simple to implement and computationally efficient, updating fast weights within a single forward pass without inference-time backpropagation. Extensive experiments show that P-TTT more effectively captures historical preference relations and outperforms state-of-the-art methods by a large margin.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑