arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.16553cs.CL

STAGE:面向多偏好大语言模型对齐的可控目标准入

STAGE: Controlled Objective Admission for Multi-Preference LLM Alignment

Yongqi Tong, Zhenyu Zhang, Ruirui Wang, Kewei Fu, Shaoqing Lin, Sijie Dong, Jiang-Ming Yang, Xin Zhang, Jianshe Li

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对多偏好LLM对齐的目标准入时机问题,提出稳定性引导的STAGE控制器,通过活动集扩展与探测排序等方法,在多偏好对齐任务中取得优于基线的自动评估结果,为RLHF提供新控制变量。

中文摘要 AI 辅助

多偏好对齐常被表述为标量化:将奖励维度组合后进行优化,但这会留下一个未明确的时间决策:每个偏好维度应在何时进入策略优化?我们提出STAGE,一种用于可控目标准入的稳定性引导活动集控制器。STAGE从一个小型活动集开始,保留已准入的目标,当奖励偏差门控显示近期偏差较低或耐心预算耗尽时进行扩展;探测阶段估计从难到易的顺序,自适应加权则强调表现不佳的活动维度。对15个训练偏好和16个保留基准列的自动评估显示,STAGE的平均值高于同时标量化及适配共享预算的基线。组件消融与扩展动态进一步支持,在该场景中,累积保留、门控准入和探测衍生的排序是有用的设计选择。这些结果将目标准入时机定位为奖励向量RLHF中的一个具体控制变量。

英文摘要

Multi-preference alignment is often framed as scalarization: combine reward dimensions, then optimize. This leaves a temporal decision underspecified: when should each preference dimension enter policy optimization? We propose \methodname, a stability-guided active-set controller for controlled objective admission. \methodname starts from a small active set, retains admitted objectives, and expands when reward-deviation gates indicate low recent deviation or a patience budget is exhausted. A probing phase estimates a hard-to-easy order, and adaptive weighting emphasizes underperforming active dimensions. Automatic evaluations with 15 training preferences and 16 held-out benchmark columns show that \methodname obtains higher averages than simultaneous scalarization and shared-budget adapted baselines. Component ablations and expansion dynamics further support cumulative retention, gated admission, and probing-derived ordering as useful design choices in this setting. These results position objective-entry timing as a concrete control variable in reward-vector RLHF.

发表机构

  • Ant International(蚂蚁国际)

机构由 AI 辅助整理,请以论文原文为准。

↑