arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

后训练中的锐化税

Sharpening Tax in Post-Training

Changdae Oh, Qi Zeng, Qi Qi, Andrey Zhmoginov, Deren Lei, Yun He, Hoang Phan, Hangoo Kang, Azalia Mirhoseini, Sharon Li

arXiv 2610.01509首次发表:更新:

发表机构

Meta Superintelligence Labs; University of Wisconsin–Madison; NYU; Stanford University(Meta超级智能实验室; 威斯康星大学麦迪逊分校; 纽约大学; 斯坦福大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出锐化税指标量化后训练导致的测试时可扩展性损失,发现预训练模型在覆盖范围上优于后训练模型,并设计PTGS采样器以降低该税。

AI 中文摘要

关于大型语言模型(LLM)强化学习(RL)后训练的一个新兴假设是,它仅仅锐化了基础模型的现有行为,以提高单次采样准确率为代价,牺牲了解决方案的覆盖范围。尽管这种权衡已在数学和编码任务中被观察到,但它未必适用于智能体任务,在这些任务中,多轮工具使用和交互可能需要后训练期间新获得的能力。我们的惊人发现是,配备轻量推理框架的预训练LLM可以作为有能力的智能体。尽管准确率(pass@1)远低于后训练模型,但在给定足够的测试时预算的情况下,它们在解决方案覆盖范围(pass@K)上往往超过后训练模型。我们进一步分析了潜在机制,并表明后训练将任务推向两个极端:要么总是解决,要么从不解决,从而以牺牲解决方案覆盖范围为代价,提高了采样效率和一致性。为了衡量这一代价,我们提出了锐化税(Sharpening Tax),这是一种诊断指标,用于量化后训练后测试时可扩展性的损失。在来自四个模型家族和三个智能体基准的14个基础/后训练模型对(共42个案例)中,该税在大多数设置中普遍存在,可以从少量回滚中估计,并与其他指标良好相关。最后,我们提出了后验温度组采样(PTGS),这是一种简单的即插即用贝叶斯采样器,根据每个提示的估计难度调整其采样温度。在两个智能体环境中应用于RL训练时,PTGS比固定温度基线支付更小的税,在重复采样下解决更多任务,同时提高单次采样准确率。

英文摘要

An emerging hypothesis about reinforcement learning (RL) post-training of large language models (LLMs) is that it merely sharpens existing behaviors of a base model, improving single-shot accuracy at the cost of solution coverage. Although this trade-off has been observed in math and coding tasks, it need not extend to agentic tasks, where multi-turn tool use and interaction may require capabilities newly acquired during post-training. Our surprising finding is that pre-trained LLMs, equipped with a light inference harness, can serve as capable agents. Despite far lower accuracy (pass@1), they often surpass their post-trained counterparts in solution coverage (pass@K) given a sufficient test-time budget. We further analyze the underlying mechanism and show that post-training pushes tasks toward two extremes, always solved or never solved, and thereby improves sampling efficiency and consistency at the cost of solution coverage. To measure this cost, we propose Sharpening Tax, a diagnostic metric that quantifies the loss in test-time scalability after post-training. Across 14 base/post-trained model pairs from four families and three agentic benchmarks (42 cases in total), the tax is prevalent in most settings, can be estimated from a few rollouts, and correlates well with other metrics. Finally, we present posterior-tempered group sampling (PTGS), a simple plug-and-play Bayesian sampler that adapts the sampling temperature per prompt to its estimated difficulty. Applied during RL training in two agentic environments, PTGS pays a smaller tax than the fixed-temperature baseline, solving more tasks under repeated sampling while also improving single-shot accuracy.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑