动态路由:大语言模型测试时多功能性的新维度
Dynamic Routing as a New Dimension for Test-time Versatility of LLMs
浏览论文内容
中文总结 AI 辅助
本文提出动态路由作为大语言模型测试时适应的新维度,与思维链互补且无需梯度更新,通过少量示范和置信度选择策略,在MMLU和ARC-AGI等任务上显著提升准确率,并指出其潜力与挑战。
中文摘要 AI 辅助
除了扩展参数和数据规模,大型语言模型目前在新问题上获得多功能性的途径仅有一个维度:它们在思维链(CoT)上花费的令牌数。我们研究动态路由程序(执行模型的部分层或迭代其中某些层)是否能开辟测试时适应的第二个维度,该维度与CoT互补且无需任何梯度更新。先前研究表明,此类程序存在,并在与其训练数据相似的问题上带来准确性和效率提升;我们提出的问题是:这些程序能否通过一种跨模型和任务迁移且无需训练的策略,仅凭少量示范(3个或10个)被快速识别。首先,我们发现,通过模型自身置信度仲裁的、依据程序分配给示范标签的概率来选择的策略,能带来一致的收益:在MMLU的49个任务上平均表现提升显著,在七个模型中的四个上提升明显,尤其是在分布外任务(如ARC-AGI)上效果最佳,其中程序使一个CoT失效的7B模型的准确率翻倍。其次,在七个后训练模型的MMLU任务上,路由在实践中与CoT互补:两者在不同查询上成功,其组合性能超过单独使用CoT。尽管有这些收益,我们的分析表明,基于置信度的选择留下了大量未开发的潜力,特别是在两个方面:(1)在揭示预训练早期已存在但后训练后更难选择的路由潜力方面,(2)在使模型对路由引入的细化调整具有鲁棒性方面,因为不成功的路由往往会使残差流偏离后续层期望的分布。综合来看,我们的结果指出动态路由是扩展现有和未来大语言模型在快速测试时适应中可塑性的一种范式。
英文摘要
Beyond scaling their parameters and data, large language models currently gain versatility on new problems along a single axis: the tokens they spend on chain-of-thought (CoT). We investigate whether dynamic routing programs, which execute a subset of the model's layers or iterate some of them, can open a second axis of test-time adaptation, complementary to CoT and free of any gradient update. Prior work showed that such programs exist and bring accuracy and efficiency gains on problems similar to those they were trained on; we ask whether they can also be identified rapidly, from a handful of demonstrations (3 or 10), by a strategy that transfers across models and tasks without training. First, we find that strategies that select programs by the probability they assign to the demonstrations' labels, arbitrated by the model's own confidence, bring consistent gains: on average over the 49 tasks of MMLU and substantially on four of seven models, and most of all on far out-of-distribution tasks such as ARC-AGI, where programs double the accuracy of a 7B model whose CoT fails. Second, on MMLU across the seven post-trained models, routing complements CoT in practice: the two succeed on different queries, and their composition exceeds CoT alone. Despite these gains, our analyses show that confidence-based selection leaves much of the potential untapped, in two places in particular: (1) in surfacing the routing potential that is already present early in pre-training but becomes harder to select after post-training, and (2) in making models robust to the refinements routing introduces, since unsuccessful routes tend to drive the residual stream out of the distribution that the following layers expect. Together, our results point to dynamic routing as a paradigm for extending the plasticity of existing and future LLMs in rapid test-time adaptation.
发表机构
- TransformersClub, Faculty of Informatics, Masaryk University(马萨里克大学信息学院 TransformersClub)
- R&D Centre for Large Language Models, National Institute of Informatics, Japan(日本国立信息学研究所大语言模型研发中心)
机构由 AI 辅助整理,请以论文原文为准。