发表机构
KAIST; Upstage; University of Toronto; Carnegie Mellon University; NVIDIA(韩国科学技术院; Upstage; 多伦多大学; 卡内基梅隆大学; 英伟达)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对不可验证强化学习中训练提示静态导致的问题,提出LLM-as-a-Tutor框架,让大语言模型从评判扩展为导师,通过对比策略展开检测无挑战性提示并添加约束,提升性能。
AI 中文摘要
不可验证指令跟随的强化学习越来越依赖大语言模型评判作为奖励信号。近期方法虽在训练中调整评判标准,但训练提示仍静态。本文提出LLM-as-a-Tutor框架,将大语言模型角色从评判扩展为导师,能检测无挑战性提示并添加原子约束,产生自校准训练信号,在基准测试中表现优于基线和先前方法。
英文摘要
Reinforcement learning (RL) for non-verifiable instruction following increasingly relies on LLM judges with prompt-specific rubrics as reward signals. While recent methods adapt these rubrics to the evolving policy during training, the training prompts themselves remain static, drawn from fixed corpora. This static approach often results in a critical misalignment between prompt difficulty and policy capability, leaving the judge unable to recover a discriminative reward signal when prompts fail to elicit quality variance among rollouts. To address this misalignment, we introduce LLM-as-a-Tutor, a framework that extends the LLM's role from judge to tutor: a single model serves as an examiner that pairwise-compares policy rollouts to detect non-challenging prompts, and as a generator that appends atomic constraints to them. This append-only design monotonically raises difficulty in step with the policy's capability, producing a self-calibrating training signal without external difficulty schedules. On three complex instruction-following benchmarks, our method consistently outperforms both policy-unaware baselines and prior policy-adaptive methods that adapt rubrics or rewrite prompts, suggesting prompt adaptation as a missing axis of policy-awareness in non-verifiable RL.