arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Hi-TTRL:利用提示调控测试时强化学习的共识

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning

Kunbin Xu, Xingzuo Li, Xuefeng Bai, Kehai Chen

arXiv 2608.03545首次发表:更新:

AI 中文总结

该研究针对测试时强化学习(TTRL)中共识强度不当导致的问题,提出 Hi-TTRL 框架,通过 MCMC 提示采样器调控共识强度,提升了 TTRL 的性能。

AI 中文摘要

测试时强化学习(TTRL)通过利用多数投票构建的伪标签更新策略,在无需标注数据的情况下提升了大语言模型的推理能力。尽管该方法有效,但多数投票分配的奖励信号对共识强度高度敏感,共识强度被定义为 rollout 组内最常见答案的出现频率。在 TTRL 中,共识强度具有双重作用:它既反映了伪标签的可靠性,也体现了优势值的分布。低共识会通过不成比例的大优势值放大不可靠伪标签的更新,而高共识则会降低奖励对比度,最终导致梯度消失。本文提出了 Hi-TTRL,一种在采样阶段利用提示调控 rollout 共识强度的测试时强化学习框架。Hi-TTRL 首先从部分 rollout 组估计共识强度;当共识强度超出目标区间时,它会调用马尔可夫链蒙特卡洛(MCMC)提示采样器,该采样器以幂变换后的前缀分布为目标,通过有限步近似采样生成 rollout 前缀作为提示。通过调整幂指数,Hi-TTRL 可生成具有增强或扁平化幂目标的提示,将 rollout 共识强度引导至目标区间。在多个数据集和主干模型上的实验表明,Hi-TTRL 始终优于标准 TTRL; ablation 实验和共识引导分析验证了自适应提示引导的共识调控的有效性。

英文摘要

Test-time reinforcement learning (TTRL) improves the reasoning capabilities of large language models without labeled data by updating the policy with pseudo-labels constructed through majority voting. While effective, the reward signal assigned from majority voting is highly sensitive to consensus strength, defined as the frequency of the most common answer within a rollout group. In TTRL, consensus strength plays a dual role: it reflects both the reliability of the pseudo-label and the distribution of advantages. Low consensus can amplify updates from unreliable pseudo-labels through disproportionately large advantages, whereas high consensus reduces reward contrast and ultimately yields vanishing gradients. In this paper, we introduce Hi-TTRL, a test-time reinforcement learning framework that utilizes hints during sampling to regulate rollout consensus strength. Hi-TTRL first estimates consensus strength from a partial rollout group. When the consensus strength falls outside a target interval, it invokes a Markov chain Monte Carlo (MCMC) hint sampler. The sampler targets the power-transformed prefix distribution and uses finite-step approximate sampling to generate rollout prefixes as hints. By tuning the power exponent, Hi-TTRL generates hints with a sharpened or flattened power target, steering rollout consensus strength toward the target interval. Experiments on multiple datasets and backbones show that Hi-TTRL consistently improves over standard TTRL, with ablations and consensus-steering analyses validating the effectiveness of adaptive hint-guided consensus regulation.

Comments15 pages, 7 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑