小模型为大模型数据控制探寻瓶颈顺序
Small Models Scout Bottleneck Order for Large-Model Data Control
AI总结:
该研究提出 LogFloor 闭环控制器,利用小模型 scout 路径的阶段排序瓶颈解决结构,在大模型数据控制中显著降低 token 成本,验证了该结构的可迁移性。
AI中文摘要:
小代理模型通常用于识别大规模训练的数据混合情况。本文探究其训练轨迹是否能揭示另一种可迁移结构:大模型解决技能瓶颈的顺序。我们提出首次通过技能训练,其中每个被监控技能都有一个目标阈值,目标是最小化达到所有阈值所需的 token 数量。我们引入 LogFloor,这是一种闭环控制器,可将每一轮训练导向当前瓶颈,生成按阶段排序的解决轨迹。在 Qwen2.5-1.5B 的五个 bAbI 技能切片上,LogFloor 平均降低了 56.2% 的 token 成本。在 70M 到 12B 的迁移中,对 70M scout 路径进行三轮回放,在所有八次目标运行中都达到了每个阈值,按配对均值计算节省了 30.9%,按汇总训练 token 计算节省了 39.4%,按源成本核算则节省了 37.6%。在 MMLU-control 上,冻结的 scout 路径在所有八次 12B 运行中均成功。将路径简化为静态边际混合或反转其阶段顺序会消除大部分收益,而仅瓶颈标签仍有部分用处。这些结果表明,按阶段排序的瓶颈解决是针对被监控技能目标训练的可迁移课程结构。
英文摘要:
Small proxy models are commonly used to identify data mixtures for larger-scale training. We ask whether their training trajectories reveal another transferable structure: the order in which larger models should resolve skill bottlenecks. We formulate first-passage skill training, where each monitored skill has a target floor and the objective is to minimize the tokens required to reach all floors. We introduce LogFloor, a closed-loop controller that directs each round toward current bottlenecks, producing phase-ordered resolution trajectories. Across five bAbI skill slices on Qwen2.5-1.5B, LogFloor reduces token cost by 56.2% on average. In 70M-to-12B transfer, three-round replay of a 70M scout path reaches every floor in all eight target runs, saving 30.9% by pair mean, 39.4% in pooled training tokens, and 37.6% under source-cost accounting. On MMLU-control, a frozen scout path succeeds across all eight 12B runs. Collapsing a path to its static marginal mixture or reversing its phase order removes most benefits, while bottleneck labels alone remain partially useful. These results identify phase-ordered bottleneck resolution as a transferable curriculum structure for monitored skill-targeted training.