在变压器推理中为潜在算法路由建立基础
Grounding latent algorithm routing in transformer reasoning
- Georgia Institute of Technology(佐治亚理工学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究变压器能否围绕不同归纳偏差族组织情节级适应,通过潜在算法路由及ROUTEBENCH基准实验,发现从零训练的密集仅解码器变压器能开发类似路由的内部变量,缩小最优路由差距,但未建立通用路由。
AI中文摘要:
上下文学习文献中的一个核心问题是变压器能否围绕不同的归纳偏差族组织情节级别的适应。我们通过潜在算法路由在受控环境中研究这个问题:求解器族偏好随潜在数据生成机制变化而提示形式固定的类似路由行为,在干扰扰动下保持稳定,并受目标激活干预选择性影响且答案质量损失不大。我们引入ROUTEBENCH,一个诊断基准,其机制不同程度地有利于全局收缩、稀疏性、鲁棒性和局部性,由类似岭回归、套索回归、Huber回归和kNN的族代表实现。在从零开始训练的44M - 612M参数的密集仅解码器变压器中,一个306M模型缩小了80.9%的最优路由差距,实现了84.1的路由F1。在自然语言渲染等等情况下效果依然显著。更强的自适应替代方案缩小了差距但在路由F1和OOD性能上仍低于306M和612M模型。探测控制和匹配激活修补控制进一步表明与路由相关的内部方向是可解码的且在求解器族一致的输出行为中起作用。这些结果提供了受控证据表明在ROUTEBENCH上训练的密集变压器可以开发类似路由的内部变量,但未在预训练语言模型或无限制自然语言推理中建立通用路由。
英文摘要:
A central question in the in-context learning literature is whether transformers can organize episode-level adaptation around different inductive-bias families. We study this question in a controlled setting through latent algorithm routing: route-like behavior in which the solver-family preference changes with the latent data-generating regime while prompt form is held fixed, remains stable under nuisance perturbations, and is selectively influenced by targeted activation interventions without large losses in answer quality. We introduce ROUTEBENCH, a diagnostic benchmark whose regimes differentially favor global shrinkage, sparsity, robustness, and locality, operationalized by ridge-like, lasso-like, Huber-like, and kNN-like family representatives. Across dense decoder-only transformers trained from scratch at 44M-612M parameters, a 306M model closes 80.9 percent of the oracle-routing gap and achieves route F1 of 84.1. The effect remains substantial under natural-language renderings, shuffled supports, lexical paraphrases, and a unified four-way routing setting. Stronger adaptive alternatives, including an input-conditioned soft mixture and an unsupervised Gumbel router, narrow the gap but remain below the 306M and 612M models on route F1 and OOD performance. Probe controls and matched activation-patching controls further show that route-relevant internal directions are decodable and functionally involved in solver-family-consistent output behavior. These results provide controlled evidence that dense transformers trained on ROUTEBENCH can develop route-like internal variables, but they do not establish universal routing in pretrained language models or unrestricted natural-language reasoning.