STAGE:面向多偏好大语言模型对齐的可控目标准入
STAGE: Controlled Objective Admission for Multi-Preference LLM Alignment
浏览论文内容
中文总结 AI 辅助
该研究针对多偏好LLM对齐的目标准入时机问题,提出稳定性引导的STAGE控制器,通过活动集扩展与探测排序等方法,在多偏好对齐任务中取得优于基线的自动评估结果,为RLHF提供新控制变量。
中文摘要 AI 辅助
多偏好对齐常被表述为标量化:将奖励维度组合后进行优化,但这会留下一个未明确的时间决策:每个偏好维度应在何时进入策略优化?我们提出STAGE,一种用于可控目标准入的稳定性引导活动集控制器。STAGE从一个小型活动集开始,保留已准入的目标,当奖励偏差门控显示近期偏差较低或耐心预算耗尽时进行扩展;探测阶段估计从难到易的顺序,自适应加权则强调表现不佳的活动维度。对15个训练偏好和16个保留基准列的自动评估显示,STAGE的平均值高于同时标量化及适配共享预算的基线。组件消融与扩展动态进一步支持,在该场景中,累积保留、门控准入和探测衍生的排序是有用的设计选择。这些结果将目标准入时机定位为奖励向量RLHF中的一个具体控制变量。
英文摘要
Multi-preference alignment is often framed as scalarization: combine reward dimensions, then optimize. This leaves a temporal decision underspecified: when should each preference dimension enter policy optimization? We propose \methodname, a stability-guided active-set controller for controlled objective admission. \methodname starts from a small active set, retains admitted objectives, and expands when reward-deviation gates indicate low recent deviation or a patience budget is exhausted. A probing phase estimates a hard-to-easy order, and adaptive weighting emphasizes underperforming active dimensions. Automatic evaluations with 15 training preferences and 16 held-out benchmark columns show that \methodname obtains higher averages than simultaneous scalarization and shared-budget adapted baselines. Component ablations and expansion dynamics further support cumulative retention, gated admission, and probing-derived ordering as useful design choices in this setting. These results position objective-entry timing as a concrete control variable in reward-vector RLHF.
发表机构
- Ant International(蚂蚁国际)
机构由 AI 辅助整理,请以论文原文为准。