Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients
近端策略优化区域:教师存在于提示中,而非梯度中
机构 * NVIDIA(英伟达)
AI总结 提出ZPPO方法,通过将教师知识注入提示而非策略梯度,解决小模型知识蒸馏中教师梯度主导和强化学习策略漂移问题,在多种规模模型上超越现有方法。
Comments Project page: https://byungkwanlee.github.io/ZPPO-page/