arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.11548cs.LG

从受控预训练混合数据到代码的条件迁移

Conditional Transfer from Controlled Pretraining Mixtures to Code

Ohad Rubin

AI总结:

该研究区分了任务的诊断性、可教性与数据源的迁移性,通过受控预训练发现合成任务族的可教性存在不对称性,且损失导向的自适应调度器会使混合数据偏离可迁移区域。

AI中文摘要:

合成任务越来越多地被用作语言模型能力的探针以及预训练数据,这两种用途通常都以损失降低为依据:损失下降被视为具有信息价值,而通过更多采样实现的更快损失降低则被作为某一任务值得采样的证据。我们区分出三种信号:当某一任务的损失跟踪全局预训练进展时,它具有诊断性;当某一任务的损失对其自身的 token 预算有响应时,它具有可教性;当包含某一数据源能提升下游目标时,该数据源具有迁移性。我们开展了受控预训练研究,其中 70% 的语料库为固定的通用 Python 代码,剩余 30% 是三个源族的单纯形混合:OpenCodeInstruct(一套精心整理的 12 个与代码相关的合成任务)、以及 15 个源自文献的探针任务。在任务预算扫描中,我们检测到 27 个任务中有 14 个具有可教性,且两个合成任务族之间存在明显的不对称性:精心整理的合成任务族为 10/12,文献探针任务族为 4/15。可教性与下游迁移性给出了不同的排序。在精心整理任务集与 OpenCodeInstruct 之间的混合单纯形边缘,经过固定微调阶段后的 HumanEval pass@20 指标,从纯精心整理数据时的 15.9,在 OpenCodeInstruct 占比 75% 的混合数据下达到观测到的最高值 22.6,随后在纯 OpenCodeInstruct 时降至 19.5。因此,精心整理的合成数据具有条件价值:它仅在以目标对齐源为主导的混合数据中占有限份额时才发挥作用。最后,基于损失的自适应调度器揭示了残余损失可降低性与下游迁移性之间的不匹配。在三次 6 万步的自由比例运行中,Ado 在最初 5 千步内将 OpenCodeInstruct 的占比降至 5% 以下,训练结束时降至 1.2%–1.4%,且其表现比匹配的固定混合对照组低 2.4–11.0 个百分点。优化近期任务损失降低会使混合数据偏离可迁移的区域。

英文摘要:

Synthetic tasks are increasingly used both as probes of language-model capability and as pretraining data. Both uses are often justified by loss reduction: falling loss is treated as informative, and faster loss reduction with more sampling as evidence that a task is worth sampling. We separate three signals. A task is diagnostic when its loss tracks global pretraining progress; it is teachable when its loss responds to its own token budget; and a data source transfers when including it improves a downstream target. We study controlled pretraining in which 70% of the corpus is fixed general Python and the remaining 30% is a simplex over three source families: OpenCodeInstruct, a curated suite of 12 code-adjacent synthetic tasks, and 15 literature-derived probe tasks. Across a task-budget sweep we detect teachability for 14 of 27 tasks, with a sharp asymmetry between the two synthetic families (10/12 curated versus 4/15 literature-derived). Teachability and downstream transfer give different rankings. On the mixture-simplex edge between the curated suite and OpenCodeInstruct, HumanEval pass@20 after a fixed fine-tuning stage rises from 15.9 at pure curated data to its highest observed value, 22.6, at a mixture that is 75% OpenCodeInstruct, then falls to 19.5 at pure OpenCodeInstruct. Curated synthetic data therefore has conditional value: it contributes as a limited share of a mixture that a target-aligned source still dominates. Finally, a loss-based adaptive scheduler exposes the mismatch between residual loss reducibility and downstream transfer. Across three 60k-step free-ratio runs, Ado drives the OpenCodeInstruct share below 5% within the first 5k steps and to 1.2--1.4% by the end of training, and underperforms its matched fixed-mixture controls by 2.4--11.0 percentage points. Optimizing near-term task-loss reduction moves the mixture away from the region that transfers.

↑