arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

策略性多样采样用于自训练

Strategically Diverse Sampling for Self-Training

Alexander Gurung, Esmeralda S. Whitammer, Mirella Lapata

arXiv 2609.31571首次发表:更新:

发表机构

University of Edinburgh(爱丁堡大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出策略性多样采样(含GROOT和言语化采样)构建自训练数据,证明方法多样性比正确性或教师规模更重要,能提升困难任务性能并增强RL与测试时扩展。

AI 中文摘要

许多大型语言模型的训练和推理方法,包括强化学习和测试时扩展,都依赖于重复采样,但只有当响应在意义上有所差异时才能获益。自训练面临同样的挑战:训练数据通常通过独立同分布采样响应并主要基于正确性进行过滤来构建,从而过度代表了模型已经偏好的策略。我们研究策略性多样性,即解决问题方法之间的实质性差异,作为构建自训练数据的替代原则。我们通过两种采样方法生成策略性多样数据:GROOT,一种新方法,构建方法的分层树并采样不同路径;以及言语化采样(VS),改编以产生非结构化的方法集合。在竞争性编程和下一章预测领域,使用策略性采样数据训练的模型在困难任务上优于独立同分布训练的对应模型,并为强化学习和测试时扩展提供了强有力的初始化。最引人注目的是,基于Qwen3-4B的策略性多样但不正确的轨迹进行自训练,优于从235B教师模型进行独立同分布蒸馏。这些结果挑战了关于什么构成有用自训练数据的普遍假设,并表明方法的多样性可能比正确性或教师规模更重要。

英文摘要

Many LLM training and inference methods, including RL and test-time scaling, depend on repeated sampling, but benefit only when the responses meaningfully differ. Self-training faces the same challenge: training data is typically constructed by sampling IID responses and filtering primarily for correctness, thereby overrepresenting strategies a model already favours. We investigate strategic diversity, or substantive variation among approaches to a problem, as an alternative principle for constructing self-training data. We generate strategically diverse data with two sampling methods: GROOT, a new method which constructs a hierarchical tree of approaches and samples distinct paths, and Verbalized Sampling (VS), adapted to produce an unstructured set of approaches. Across competitive programming and Next-Chapter Prediction domains, models trained on strategically sampled data outperform IID-trained counterparts on difficult tasks and provide strong initializations for RL and test-time scaling. Most strikingly, self-training on strategically diverse but incorrect traces from Qwen3-4B outperforms IID distillation from a 235B teacher. These results challenge prevailing assumptions about what makes useful self-training data and show that diversity of approaches can matter more than correctness or teacher scale.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑