FlowBalance:基于验证器的在线策略推理经验自改进方法
FlowBalance: Verifier-Grounded Self-Improvement from On-Policy Reasoning Experience
- Tencent HY LLM Frontier(腾讯HY大模型前沿(机构))
- University of Pennsylvania(宾夕法尼亚大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
FlowBalance是基于验证器的在线策略推理经验自改进方法,通过校准轨迹级自指导分数实现结果校准的自指导,在数学推理任务中其性能、训练速度及稳定性均优于FlowRL,且策略多样性更高。
AI中文摘要:
推理模型可从自身在线策略经验中改进,但该内部循环存在脆弱性:终端验证器提供可靠但稀疏的监督,而密集的同模型指导可能强化错误信心或使学习过度集中于狭窄的解模式。我们提出FlowBalance,一种基于验证器的自改进方法,学习完整响应的归一化分布。对于每个在线策略轨迹,训练时冻结的同策略视图利用特权上下文生成令牌级对数概率增益,聚合为轨迹级自指导分数。FlowBalance用验证器导出的组优势校准该分数:保留正优势轨迹的指导,反转负优势轨迹的指导,当回退组无结果偏好时禁用指导。所得能量指数加权参考策略,而轮廓轨迹平衡用每个回退组的一个对数配分估计拟合归一化目标,通过轨迹平衡实现结果校准的自指导,无需单独的令牌级模仿损失。我们的分析确立了组内对比保留、最小变化反向KL表征、目标奖励的单调验证器控制,以及对被拒响应上假阳性自指导的精确校正。在数学推理任务中,FlowBalance在Qwen3-4B和Qwen3-8B上的平均性能均优于FlowRL,同时提升了训练速度与稳定性,避免了直接OPSD的响应长度崩溃,并在受控AIME24诊断中表现出更高的正确策略多样性。
英文摘要:
A reasoning model can improve from its own on-policy experience, but this inner loop is fragile: terminal verifiers provide reliable yet sparse supervision, while dense same-model guidance can reinforce false confidence or overconcentrate learning on a narrow solution mode. We introduce FlowBalance, a verifier-grounded self-improvement method that learns a normalized distribution over complete responses. For each on-policy trajectory, a frozen training-time view of the same policy uses privileged context to produce token-level log-probability gains, which are aggregated into a trajectory-level self-guidance score. FlowBalance calibrates this score with the verifier-derived group advantage: guidance is retained on positive-advantage trajectories, reversed on negative-advantage trajectories, and disabled when the rollout group provides no outcome preference. The resulting energy exponentially reweights a reference policy, and profiled trajectory balance fits the normalized target with one log-partition estimate per rollout group. This realizes outcome-calibrated self-guidance via trajectory balance, without a separate token-level imitation loss. Our analysis establishes within-group contrast preservation, a minimum-change reverse-KL characterization, monotonic verifier control of target reward, and an exact correction against false-positive self-guidance on rejected responses. On mathematical reasoning, FlowBalance improves average performance over FlowRL on both Qwen3-4B and Qwen3-8B, while also improving training speed and stability, avoiding direct OPSD's response-length collapse, and exhibiting higher correct-strategy diversity in a controlled AIME24 diagnostic.