LocusRL:诊断大型语言模型在竞争性游戏中的奖励与策略干预
LocusRL: Diagnosing LLM Reward and Policy Interventions in Competitive Games
浏览论文内容
中文总结 AI 辅助
LocusRL是一个诊断框架,通过追踪奖励与策略干预在竞争游戏中的实际影响,揭示相似回报下隐藏的学习机制,并提供可验证的修正,从而将聚合结果转化为可操作诊断。
中文摘要 AI 辅助
大型语言模型可以通过奖励设计和动作选择来干预强化学习,然而聚合性能对这些干预的实际作用提供了不完整的描述。相似的回报可能掩盖不同的学习机制,而看似合理的奖励可能诱发不良行为。我们引入了LocusRL,一个诊断框架,它将受控的奖励-策略比较与对奖励判断、信号传递、优化目标和执行动作的审计联系起来。该框架将性能差异追溯到可测试的解释,并通过可执行规则和反事实重放来检查针对性的修正。在覆盖十个四子棋训练种子的两个评估批次中,我们发现了干预效果中依赖种子的反转,并展示了追踪实际更新如何改变其解释:历史Qwen训练通过奖励加权的教师动作似然进行操作。一个独立的匹配三种子奖励方向实验区分了对学习信号的敏感性与其实用性。在终端奖励保持不变的情况下,符号反转的密集神谕产生了2.8%的聚合胜率,而仅终端训练为57.2%,正密集神谕为46.7%。因此,一个奖励可以在不提高性能的情况下强烈影响学习。在决策层面,反事实重放验证了对诊断出的动作错误的获胜修正。在Leduc中的补充实验和Goofspiel中的奖励验证研究将分析扩展到不完全信息设置,揭示了参考标签定义和验证数据暴露如何影响干预评估。总之,这些发现表明,评估LLM干预需要追踪其输出如何成为学习信号和动作。LocusRL将聚合结果转化为可操作的诊断和可验证的修正。
英文摘要
Large language models can intervene in reinforcement learning through both reward design and action selection, yet aggregate performance offers an incomplete account of what these interventions actually do. Similar returns can conceal different learning mechanisms, while plausible rewards can induce undesirable behavior. We introduce LocusRL, a diagnostic framework that connects controlled reward-policy comparisons with audits of reward judgments, signal delivery, optimization objectives, and executed actions. The framework traces performance differences to testable explanations and checks targeted corrections through executable rules and counterfactual replay. Across two evaluation batches covering ten Connect Four training seeds, we uncover seed-dependent reversals in intervention effects and show how tracing actual updates changes their interpretation: historical Qwen training operates through reward-weighted teacher-action likelihood. A separate matched three-seed reward-direction experiment distinguishes sensitivity to a learning signal from its usefulness. With terminal rewards held fixed, a sign-reversed dense oracle yields a 2.8% aggregate win rate, compared with 57.2% for terminal-only training and 46.7% for the positive dense oracle. Thus, a reward can strongly influence learning without improving performance. At the decision level, counterfactual replay verifies a winning correction to a diagnosed action error. Complementary experiments in Leduc and reward-validation studies in Goofspiel extend the analysis to imperfect-information settings, revealing how reference-label definitions and validation-data exposure affect intervention assessment. Together, these findings show why evaluating LLM interventions requires tracing how their outputs become learning signals and actions. LocusRL turns aggregate outcomes into actionable diagnoses and verifiable corrections.
发表机构
- Harbin Institute of Technology(哈尔滨工业大学)
机构由 AI 辅助整理,请以论文原文为准。