幻影增益:针对可测量零假设的自我改进审计
Phantom Gains: Auditing Self-Improvement Against a Measured Null
- University College Dublin(都柏林大学学院)
- Georgia Institute of Technology(佐治亚理工学院)
- Dalian University of Technology(大连理工大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究针对Qwen3-8B的LoRA自训练开展审计,发现测量伪影会反转改进结论,提出用逐问题精确检验替代自然阈值修复,还揭示外部蒸馏与自训练的不同效果及自训练的破坏作用。
AI中文摘要:
语言模型是否实现了自我改进,如今越来越多地依据其在单个问题上的得失而非平均准确率来判断。追踪这些得失变化需要对两个存在噪声的估计值做差分处理,这使得它们极易受到测量伪影的影响。我们针对Qwen3-8B开展了三轮秩为32的LoRA自训练,并将其与通过相同流程的冻结对照组进行对比审计,识别出7种测量失败情况,其中每一种情况在缺少对照组时都会反转已报道的发现,且其中数种属于标准做法。基于单次贪心解码构建的账本会在未训练模型上生成能力变化,这在很大程度上是推理批处理的伪影;用于区分能力获取与能力强化的扩展统计量为该模型赋予了0.280的比率。自然阈值修复无法通过复现验证:在冻结对比中估算时,这类设计已包含的零假设不会为零。我们将其替换为针对合并基线的逐问题精确检验,控制错误发现率,该检验在任何保留的复现样本上均未检测到任何内容,且在多重检验规则、错误率和池大小下保持不变。应用于在流、规模和评估上匹配的一组任务时,审计发现外部蒸馏改进了基础模型很少触及的问题,而三种形式的自训练则没有;回归分析将这种不对称性拒绝为蒸馏整体增益更大的副产品(p < 10^-8)。在基础模型从未触及的小得多的问题集上,证据不确定,而自训练以远高于测量下限的速率破坏了基线时已解决的问题。因此,过渡级审计要求对其报告的每个统计量都有一个单独测量的零假设:这类零假设无需新实验,可从多臂研究已有的基线复现样本中构建,尽管样本数量并非如多数研究那般少。
英文摘要:
Whether a language model has improved itself is increasingly judged not by mean accuracy but by which individual problems it gains and loses. Tracking these transitions means differencing two noisy estimates, leaving them vulnerable to measurement artifacts. Auditing three rounds of rank-$32$ LoRA self-training on Qwen3-8B against a frozen control pushed through the identical pipeline, we identify seven measurement failures, each of which inverts a reported finding when its control is absent. Several are standard practice. A ledger built on a single greedy decode manufactures capability changes on an untrained model, largely an artifact of inference batching; the expansion statistic separating acquisition from sharpening assigns that same model a rate of $0.280$. The natural threshold repair does not survive replication: estimated across the frozen comparisons such a design already contains, its null stays non-zero. We replace it with a per-problem exact test against a pooled baseline under false-discovery-rate control, which detects nothing on any held-out replicate and is unchanged under the multiple-testing rule, error rate and pool size. Applied to a ladder of arms matched in stream, volume and evaluation, the audit finds that external distillation improves problems the base model rarely reaches while three forms of self-training do not; a regression rejects this asymmetry as a by-product of distillation's larger overall gain ($p < 10^{-8}$). On the far smaller set of problems the base model never reaches, the evidence is inconclusive, while self-training corrupts problems solved at baseline at rates well above the measured floor. Transition-level auditing therefore requires a separately measured null for every statistic it reports: nulls that cost no new experiments, built from baseline replicates a multi-arm study already owns, though not from as few as most possess.