超越奖励黑客:分层人形学习流水线四层代理散度
Beyond Reward Hacking: Proxy Divergence Across Four Layers of a Staged Humanoid Learning Pipeline
- London, United Kingdom(伦敦,英国)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文论证强化学习流水线中奖励、课程门、评估统计和参考运动均为代理,各有散度机制,并提出替代方案,通过人形机器人实验验证。
AI中文摘要:
腿式机器人的强化学习(RL)流水线由代理(proxy)组装而成。奖励代表预期行为,课程门代表能力,评估统计代表鲁棒性,参考运动代表可实现技能。传统观点仅将第一个视为优化目标,因此将规范失败(奖励黑客)仅归因于奖励。我认为所有四个在相同形式意义上都是代理,每个都有其特征性的散度机制,并且每个都允许通过重新表述来闭合。对于每一层,我陈述传统表述,推导其偏离目标的条件,并给出替代方案:在二次核平坦处使用一阶(L1)成本,在课程门平均处使用峰值和结果统计,门可达性和信息检查,将课程状态视为模型的一部分,可行性优先的参考设计及残差前馈,以及保持函数不变的输入扩展,使一个策略能够增长而非重新训练。这些论点通过一个模拟1.91米人形机器人的PPO策略的连续谱系测量加以说明,该策略在单个笔记本电脑GPU上经过四个阶段和13,500次迭代训练而成。其中,基于平均误差构建的课程门在每次检查时以其速率限制推进,而其所门控的技能却缺失;一个批量推挤测试的同步重置混淆了步态相位,将0.5米/秒的推挤评为比2.0米/秒的更危险。
英文摘要:
A reinforcement-learning (RL) pipeline for a legged robot is assembled from proxies. A reward stands in for intended behaviour, a curriculum gate stands in for competence, an evaluation statistic stands in for robustness, and a reference motion stands in for an achievable skill. The traditional view treats only the first of these as optimised against, and so locates specification failure (reward hacking) in the reward alone. I argue that all four are proxies in the same formal sense, that each has a characteristic divergence mechanism, and that each admits a reformulation that closes it. For every layer I state the traditional formulation, derive the condition under which it diverges from its target, and give the alternative: first-order (L1) costs where quadratic kernels are flat, peak and outcome statistics where curriculum gates average, gate reachability and information checks, deterministic and phase-desynchronised evaluation, curriculum state treated as part of the model, feasibility-first reference design with residual feed-forward, and function-preserving input widening that lets one policy grow instead of being retrained. The arguments are illustrated by measurements from one continuous lineage of a PPO policy for a simulated 1.91 m humanoid, grown over four stages and 13,500 iterations on a single laptop GPU. Among them, a curriculum gate built on averaged error advanced at its rate limit on every check while the skill it gated was absent, and a batched push test whose synchronised resets aliased the gait phase ranked a 0.5 m/s push as more dangerous than a 2.0 m/s one.