Ratio-Variance Regularized Policy Optimization
比率方差正则化策略优化
机构 * Department of Foundation Model, 2012 Labs, Huawei(华为基础模型部门,2012实验室) ; Shanghai Research Institute for Intelligent Autonomous Systems, Tongji University(上海智能自主系统研究院,同济大学) ; Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系) ; College of Intelligence and Computing, Tianjin University(天津大学智能与计算学院)
专题命中 数学推理 :reasoning(abstract);分类 cs.AI、cs.LG
AI总结 提出R²VPO方法,通过约束策略比率方差作为信任区域的局部近似,替代启发式裁剪,在LLM和机器人控制任务中提升性能与样本效率。