LambdaPO: A Lambda Style Policy Optimization for Reasoning Language Models
LambdaPO: 一种用于推理语言模型的Lambda风格策略优化
专题命中 安全训练 :alignment(abstract);分类 cs.CL
AI总结 针对GRPO因使用群体均值作为基线而丢失细粒度偏好信息的问题,提出LambdaPO方法,通过将优势估计分解为成对偏好结构并引入语义密度奖励,从群体轨迹中挖掘更细粒度的优化信号,提升推理性能。
Comments arXiv admin comment: This version has been removed by arXiv administrators as the submitter did not have the rights to agree to the license at the time of submission. Author list and submitter name redacted due to disputed authorship