针对不可验证生成的演员条件评论家的协同进化
Co-Evolving Actor-Conditioned Critics for Non-Verifiable Generation
浏览论文内容
中文总结 AI 辅助
针对不可验证生成的批评监督问题,引入TAIScore奖励和GRPO训练的演员定制评论家,构建评论家-演员协同进化循环,其性能优于现有方法,证明批评监督应随演员变化调整。
中文摘要 AI 辅助
自然语言批评为缺乏确定性验证器的不可验证生成提供了超出标量奖励的监督。在批评引导的优化过程中,评论家对初始响应提供反馈,演员对其进行修订。然而,最终修订的质量无法揭示批评是否真正有用:有能力的演员可能不遵循反馈也能改进,而有效的反馈若演员无法执行则会失效。我们将批评构建为演员条件的修订指导,其有用性取决于反馈是否帮助目标演员解决预期的弱点。我们引入TAIScore(针对性可操作改进分数),这是一种奖励,共同评估指令、初始响应、批评和修订,评估批评是否针对真实弱点、演员是否遵循批评以及预期方面是否得到改进。我们使用该奖励通过GRPO训练针对演员定制的评论家,并利用批评引导的优化为演员构建DPO偏好对,形成评论家-演员协同进化的循环,其中评论家会适应演员不断变化的能力。实验表明,用TAIScore训练的8B评论家,其性能优于零样本120B评论家以及仅用结果或仅用批评奖励信号训练的评论家。评论家与演员的协同进化进一步提升了性能,表明有效的批评监督应随演员的变化而调整。
英文摘要
Natural-language critiques provide supervision beyond scalar rewards for non-verifiable generation, which lacks deterministic verifiers. In critique-guided refinement, a critic gives feedback on an initial response and an actor revises it. However, final revision quality does not reveal whether the critique was actually useful: a capable actor may improve without following the feedback, while valid feedback may fail if the actor cannot execute it. We frame critique as actor-conditioned revision guidance, where usefulness depends on whether the feedback helps the target actor address the intended weakness. We introduce TAIScore (Targeted Actionable Improvement Score), a reward that evaluates the instruction, initial response, critique, and revision together, assessing whether the critique targets a real weakness, whether the actor follows it, and whether the intended aspect improves. We use this reward to train an actor-tailored critic with GRPO, and use critique-guided refinements to construct DPO preference pairs for the actor, forming a co-evolving critic-actor loop where the critic adapts to the actor's changing capability. Experiments show that an 8B critic trained with TAIScore outperforms both a zero-shot 120B critic and critics trained with outcome-only or critique-only reward signals. Co-evolving the critic and actor further improves performance, suggesting that effective critique supervision should adapt as the actor changes.
发表机构
- University of Michigan(密歇根大学)
- LG AI Research(LG AI研究院)
- University of Illinois at Chicago(芝加哥大学伊利诺伊分校)
机构由 AI 辅助整理,请以论文原文为准。