arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.27505cs.CLcs.AI

面向语言模型的准则引导强化学习综述

A Survey on Rubric-Guided Reinforcement Learning for Language Models

Zifei Shan, Fangning Shao

首次发表
浏览论文内容

中文总结 AI 辅助

该综述针对传统RLHF的缺陷,提出贝叶斯框架并沿先验-后验轴分类准则引导强化学习,分析语言学相关对齐问题,明确未来研究方向。

中文摘要 AI 辅助

从人类反馈的强化学习(RLHF)已成为使大型语言模型(LLM)与人类偏好对齐的主流范式。然而,传统RLHF依赖标量奖励信号,缺乏可解释性,且无法捕捉响应质量的多面性。准则引导强化学习通过引入结构化、可解释的评估准则(即rubrics)作为奖励设计、反馈生成和策略优化的核心,解决了这些局限。本综述中,我们介绍了一种贝叶斯框架,将constitution定义为评估准则上的先验分布P(R),将rubrics定义为条件实例R_x ~ P(R|x)。在此统一视角下,我们沿先验-后验轴对准则引导RL进行分类,涵盖宪法AI、实例特定准则、过程级监督、自演化准则及其智能体(agentic)和多模态扩展。此外,由于rubrics是自然语言产物,我们对粒度权衡、语义漂移和语言奖励黑客如何影响对齐可靠性进行了语言学分析,确定了未来研究的关键开放问题。

英文摘要

Reinforcement learning from human feedback (RLHF) has become the dominant paradigm for aligning large language models (LLMs) with human preferences. However, traditional RLHF relies on scalar reward signals that lack interpretability and fail to capture the multifaceted nature of response quality. Rubric-guided reinforcement learning addresses these limitations by introducing structured, interpretable evaluation criteria, or rubrics, as the backbone of reward design, feedback generation, and policy optimization. In this survey, we introduce a Bayesian framework that defines constitutions as prior distributions $P(R)$ over evaluation criteria and rubrics as conditional instantiations $R_x \sim P(R|x)$. Under this unified view, we present a taxonomy of rubric-guided RL along the prior-posterior axis, covering constitutional AI, instance-specific rubrics, process-level supervision, self-evolving rubrics, and their agentic and multimodal extensions. Furthermore, as rubrics are natural-language artifacts, we present a linguistic analysis of how granularity trade-offs, semantic drift, and linguistic reward hacking impact alignment reliability, identifying key open problems for future research.

发表机构

  • WeChat, Tencent(腾讯微信)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑