arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

模型的模型:何时生成专家模型优于注意力、适配或微调?

Model of Models: When Does Emitting a Specialist Beat Attending, Adapting, or Tuning?

John C. Howell

arXiv 2608.21386首次发表:更新:

发表机构

Nineonefour(Nineonefour)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究对比零样本、上下文注意力等四种模型专业化机制,发现生成专家模型在匹配质量下成本更低,但在高维序列建模中不及上下文注意力,还提出可证伪论点界定各机制适用场景。

AI 中文摘要

给定一个由少量示例描述的任务,模型应如何针对该任务进行专业化?现有四种机制——零样本、上下文注意力、测试时梯度适配,以及从超网络生成专家权重——但最后一种机制的适用场景鲜有研究。我们在涵盖回归、生成、语言建模、强化学习、临床及基因组分类的六项任务中开展相同的四向对比,固定专家模型、上下文及(在可行情况下)训练预算。生成机制最显著的优势在于匹配质量下的成本:它在临床少样本分类中与最先进的摊销表格模型(TabPFN)表现相当,同时生成可复用的专家模型,而非针对每个查询重新关注支持集;且以每个实例132个浮点数的程序达到噪声水平的形状生成。在少样本正弦回归中,它在零测试时梯度步骤下比MAML低2-3个数量级——当训练预算均等时,该差距缩小至约30倍但仍存在。生成机制在高维序列建模中无法胜过上下文注意力:在匹配预训练预算下,单次适配器仅恢复上下文增益的一小部分(500万参数时为14.0±0.9%,1500万参数时为11.2±0.5%),LoRA秩扫描显示该不足是部分容量限制——捕获率随秩从5%升至21%,但远低于完全恢复。机制消融实验证实生成的专家模型确实是任务条件化的,而非记忆的先验;更具推测性的是,生成的专家模型在权重空间中可组合——对两个专家模型进行插值可追踪其功能的对应混合。我们提出一个可证伪的论点,通过每项任务的分辨率度量进行操作化,界定每种条件机制应被偏好的场景。

英文摘要

Given a task described by a few examples, how should a model be specialized to it? Four mechanisms are available -- zero-shot, in-context attention, test-time gradient adaptation, and emitting specialist weights from a hypernetwork -- yet the operating regime of the last is rarely mapped. We run the identical four-way comparison across six tasks spanning regression, generation, language modeling, reinforcement learning, and clinical and genomic classification, holding the specialist, the context, and (where we can) the training budget fixed. The clearest wins for emission are about cost at matched quality: it ties the state-of-the-art amortized tabular model (TabPFN) on clinical few-shot classification while emitting a reusable specialist instead of re-attending the support set per query, and reaches noise-floor shape generation with a $132$-float per-instance program. On few-shot sinusoid regression it is $2$--$3$ orders of magnitude below MAML at zero test-time gradient steps -- a margin that narrows to $\sim$$30\times$ but persists once training budgets are equalized. Emission cannot match in-context attention on high-dimensional sequence modeling: under matched-budget pre-training a one-pass adapter recovers only a minority of the in-context gain ($14.0\pm0.9\%$ at $5$M, $11.2\pm0.5\%$ at $15$M), and a LoRA-rank sweep shows this shortfall is a partial capacity limit -- capture climbs from $5\%$ to $21\%$ as rank grows but plateaus far below full recovery. Mechanism ablations confirm the emitted specialist is genuinely task-conditioned, not a memorized prior; and, more speculatively, emitted specialists compose in weight space -- interpolating two of them tracks the corresponding blend of their functions. We close with a falsifiable thesis, operationalized through a per-task resolution measure, bounding when each conditioning mechanism should be preferred.

Comments14 pages, 7 sections

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑