扩散奖励模型
Diffusion Reward Models
浏览论文内容
中文总结 AI 辅助
针对奖励模型忽略人类偏好多模态性的问题,提出扩散奖励模型DRM,以条件密度估计替代点估计,在五个基准上达到或超越基线,并提升下游RLHF策略性能。
中文摘要 AI 辅助
奖励模型支撑着大型语言模型的对齐,然而主流设计将每个提示-响应对简化为一个点估计或来自固定参数族的一个分布。这与人类偏好本质上的多模态特性相悖:同一响应可以以多种合理方式被评判,且没有单一分布族能覆盖所有方式。为了更好地拟合这一结构,我们引入了DRM,一种扩散奖励模型,将奖励建模重新定义为对$p(\mathbf{r}\mid x,y)$的条件密度估计。在冻结的LLM编码器条件下,一个轻量级扩散Transformer将高斯噪声去噪为奖励向量,对输出分布不作参数假设,并自然地表征其多模态结构。单一架构同时处理多属性回归和成对偏好数据,在推理时,$N$个样本构成一个经验奖励分布,可聚合成标量、方差或分位数。在五个基准上,DRM在匹配数据和骨干网络下达到或超越基线,尽管训练规模适中,仍与更大的判别式、分布式和生成式奖励模型保持竞争力,并在传统头部坍缩为点的场景中恢复了多模态奖励结构。不确定性感知的弃权(不执行)和置信下界(LCB)聚合进一步表明,DRM能利用超越标量奖励的分布信息来改进奖励模型决策。下游RLHF实验还表明,将DRM用作训练时奖励可提升策略性能,直接验证了基于扩散的奖励建模对RLHF训练的实用价值。
英文摘要
Reward models underpin the alignment of large language models, yet the dominant designs reduce each prompt--response pair to a point estimate or to a distribution from a fixed parametric family. This is at odds with human preference, which is inherently multimodal: the same response can be reasonably judged in many ways, and no single family covers all of them. To better fit this structure, we introduce DRM, a Diffusion Reward Model that recasts reward modeling as conditional density estimation over $p(\mathbf{r}\mid x,y)$. Conditioned on a frozen LLM encoder, a lightweight Diffusion Transformer denoises Gaussian noise into a reward vector, placing no parametric assumption on the output distribution and naturally representing its multimodal structure. A single architecture handles both multi-attribute regression and pairwise preference data, and at inference $N$ samples form an empirical reward distribution that can be aggregated into a scalar, a variance, or quantiles. Across five benchmarks, DRM matches or surpasses baselines under matched data and backbone, stays competitive with much larger discriminative, distributional, and generative RMs despite its modest training scale, and recovers multimodal reward structure where conventional heads collapse to a point. Uncertainty-aware rejection and lower-confidence-bound (LCB) aggregation further demonstrate that DRM can exploit distributional information beyond a scalar reward to improve reward-model decisions. Downstream RLHF experiments additionally show that using DRM as the training-time reward leads to improved policy performance, directly validating the practical benefit of diffusion-based reward modeling for RLHF training.
发表机构
- Tsinghua University(清华大学)
- The Chinese University of Hong Kong(香港中文大学)
- University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
机构由 AI 辅助整理,请以论文原文为准。