arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

先锐化再适应:测试时强化学习的数据无关入口状态锐化

Sharpen Before You Adapt: Data-Free Entry-State Sharpening for Test-Time Reinforcement Learning

Zhanming Zhang, Vinoth Selvendran

arXiv 2610.00903首次发表:更新:

AI 中文总结

提出在测试时强化学习前进行数据无关的入口状态锐化,以入口熵为控制信号,实验表明可显著提升TTRL的适配效率与最终性能。

AI 中文摘要

测试时强化学习(TTRL)利用从自身样本中获得的监督信号,在未标注的测试问题上对语言模型进行适配。这使得检查点的“入口状态”至关重要:一个弥散的策略会提供噪声更大的自监督信号,并可能在有限的适配预算中,仅仅为了集中概率质量而消耗大量资源,然后才能可靠地表达其已具备的能力。我们提出“入口状态锐化”:在TTRL之前使用数据无关训练,将通用检查点准备到一种状态,使后续的无标签适配能够更高效地利用它。这一思想并不局限于某一种训练方案;不同的数据无关目标可以将同一基础模型移动到不同的入口状态。在从Qwen3-4B派生的五个数据无关检查点上,采用相同的15步TTRL协议进行评估,入口策略熵与端点转换效率(一种可靠性到可达性的度量)之间存在强秩相关(Spearman ρ=-0.90;在控制入口可达性后ρ=-0.99)。各目标之间的对比十分显著:R-Zero保持弥散状态(3.39纳特),在MATH、GPQA和AMC上的6/6次匹配比较中均低于未调优的基础模型;而SPIRAL达到0.07纳特,并在MATH和GPQA上取得了最高的TTRL后准确率,尽管其自博弈阶段未使用任何数学训练数据。一项域内无标签自蒸馏干预进一步表明,入口状态可以被有意地锐化。这些结果促使我们将检查点准备视为一个“状态控制问题”:使用数据无关训练来提高TTRL就绪度,以入口熵作为无标签控制信号,以可达能力作为约束。

英文摘要

Test-time reinforcement learning (TTRL) adapts language models on unlabeled test problems using supervision derived from their own samples. This makes the checkpoint's \emph{entry state} consequential: a diffuse policy provides noisier self-supervision and may spend much of a limited adaptation budget merely concentrating probability mass before reliably expressing capability it already possesses. We propose \textbf{entry-state sharpening}: use data-free training \emph{before} TTRL to prepare a general-purpose checkpoint in a state that subsequent label-free adaptation can exploit more efficiently. The idea is not tied to one training recipe; different data-free objectives can move the same base model to different entry states. Across five data-free checkpoints derived from Qwen3-4B and evaluated under an identical 15-step TTRL protocol, entry policy entropy strongly rank-orders endpoint conversion efficiency, a reliability-to-reachability measure (Spearman $ρ=-0.90$; $ρ=-0.99$ after controlling for entry reachability). The contrast across objectives is striking: R-Zero remains diffuse at $3.39$ nats and finishes below the untuned base in 6/6 matched comparisons across MATH, GPQA, and AMC, whereas SPIRAL reaches $0.07$ nats and achieves the highest post-TTRL accuracy on MATH and GPQA despite its self-play stage using no math training data. An in-domain label-free self-distillation intervention further shows that the entry state can be deliberately sharpened. These results motivate treating checkpoint preparation as a \emph{state-control problem}: use data-free training to improve TTRL readiness, with entry entropy as a label-free control signal and reachable capability as the constraint.

Comments10 pages, 2 figures, 3 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑