arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TTPO:测试时策略优化

TTPO: Test-Time Policy Optimization

Aozhe Wang, Zhengxi Lu, Jianze Wang, Shangke Lv, Ying Liu, Weiming Lu, Jun Xiao, Yueting Zhuang, Hua Yang, Qianglong Chen, Yongliang Shen

arXiv 2608.27448首次发表:更新:

发表机构

Zhejiang University; Alibaba Group(浙江大学; 阿里巴巴集团)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对大语言模型测试时训练的伪标签缺陷,提出 TTPO 方法,通过不对称目标函数与 token 级选择优化,在无标签下实现了与标签监督 OPSD 相当的性能,提升了模型数学推理能力并增强了跨任务泛化性。

AI 中文摘要

近期,强化学习(RL)、在线自蒸馏(OPSD)等主流后训练方法推动了大语言模型数学推理能力的快速提升,但这类方法依赖真实标签,无法实现测试时训练(TTT)。用多数投票生成的伪标签替代真实标签是自然的替代方案,却存在缺陷:错误的投票会破坏教师模型并误导每个 token。我们发现该失效模式具有不对称性:无论投票本身是否正确,与伪标签不一致的 rollout 通常是错误的。基于此,我们提出测试时策略优化(TTPO),这是一种不对称目标函数,通过 OPSD 提炼一致的 rollout,同时用分组强化学习(Grouped RL)惩罚不一致的 rollout。token 级别的选择进一步优化了两个分支:蒸馏过程降低已收敛位置的权重,而 RL 仅惩罚置信度高的错误。即使伪标签频繁出错,这两种更新仍能保持合理性,且随着模型性能提升,多数投票路由能提供更紧密的自监督。在无需任何标签的情况下,TTPO 在五个竞赛级基准上达到了标签监督的 OPSD 的性能,将 Qwen3-1.7B 的 TTT 准确率从 38.0% 提升至 45.2%,在无思考任务上实现了 +25.2% 至 +36.4% 的提升,并展现出强大的跨任务泛化能力。

英文摘要

Recent prominent post-training methods, such as Reinforcement Learning (RL) and On-Policy Self-Distillation (OPSD), have driven rapid progress in mathematical reasoning for large language models, yet their reliance on ground-truth labels precludes test-time training (TTT). Replacing ground truth with majority-vote pseudo-labels is a natural alternative, yet it is fragile: an incorrect vote corrupts the teacher and misleads every token. We observe that this failure mode is asymmetric: rollouts that disagree with the pseudo-label are typically wrong regardless of whether the vote itself is correct. Building on this observation, we propose Test-Time Policy Optimization (TTPO), an asymmetric objective that distills agreeing rollouts via OPSD and penalizes disagreeing rollouts with Grouped RL. Token-level selection further refines both branches: distillation down-weights already-converged positions, while RL penalizes only confident errors. Both updates remain well-grounded even under frequent pseudo-label errors, and majority-vote routing yields tighter self-supervision as the model improves. Without any labels, TTPO matches label-supervised OPSD on five competition-level benchmarks, raises Qwen3-1.7B from 38.0% to 45.2% in TTT, yields +25.2% to +36.4% without thinking, and shows strong cross-task generalization.

CommentsProject Page: https://zju-real.github.io/TTPO Code: https://github.com/ZJU-REAL/TTPO

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑