arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.18587cs.LGcs.AI

无标签引导:将测试时强化学习压缩到仅偏置子空间

Label-free steering: Compressing test-time reinforcement learning into bias-only subspaces

Naveen Vakada, Mingyuan Li, Shaoxiong Ji

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出无标签仅偏置测试时强化学习,仅优化约10万偏置参数,在MATH-500上达76.67%准确率,且可迁移至其他任务,证明极小参数空间可实现有效适应。

中文摘要 AI 辅助

测试时强化学习(TTRL)使模型无需依赖有标签的训练数据即可提升推理能力,但现有方法通常优化模型参数中的很大一部分。这引出一个自然的问题:当奖励信号和优化空间都受到严格限制时,有效的测试时适应能否出现?我们通过无标签仅偏置TTRL回答了这个问题,该方法使用多数投票伪标签作为奖励,仅优化约10万偏置参数,同时保持预训练主干网络冻结。在MATH-500上,我们的方法达到了76.67%的准确率,略高于我们自己的有标签偏置引导复现结果,同时比全参数TTRL少优化了76,000倍的参数。相同的训练流程在视觉-语言和音频推理任务(包括MathVista、AI2D、LogicVista和MMAU)上也提升了性能。我们进一步表明,学习到的引导向量可迁移到4,500个留出的MATH问题,表明这种适应不仅限于测试时优化所用的问题。最后,我们分析了这种高度受限的适应为何有效,表明多数投票的可靠性随rollout共识的提高而提升,且具有更大可访问梯度能量的偏置子空间表现出更强的下游可训练性。这些结果表明,通过仅优化一个微小的仅偏置子空间并使用完全无标签的奖励,可以产生显著的测试时适应能力。

英文摘要

Test-time reinforcement learning (TTRL) enables models to improve their reasoning without relying on labeled training data, but existing approaches typically optimize a large fraction of the model parameters. This raises a natural question: can effective test-time adaptation emerge when both the reward signal and the optimization space are severely restricted? We answer this question with label-free bias-only TTRL, which uses majority-vote pseudolabels as rewards and optimizes only ~100K bias parameters while keeping the pretrained backbone frozen. On MATH-500, our approach reaches 76.67% accuracy with Qwen2.5-7B, slightly exceeding our own labeled bias-steering reproduction while optimizing 76,000x fewer parameters than full-parameter TTRL. The same training procedure improves performance across vision-language and audio reasoning tasks, including MathVista, AI2D, LogicVista, and MMAU. We further show that the learned steering vectors transfer to 4,500 held-out MATH problems, indicating that the adaptation is not limited to the problems used during test-time optimization. Finally, we analyze why this highly restricted adaptation can work, showing that majority-vote reliability improves with rollout consensus and that bias subspaces with greater accessible gradient energy exhibit stronger downstream trainability. These results demonstrate that substantial test-time adaptation can emerge from optimizing a tiny bias-only subspace using entirely label-free rewards.

发表机构

  • University of Turku(图尔库大学)
  • ELLIS Institute Finland(芬兰ELLIS研究所)

机构由 AI 辅助整理,请以论文原文为准。

↑