发表机构
Georgia Institute of Technology(佐治亚理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
预训练Transformer仅用少量深度处理上下文引用,在早期层添加任务训练的秩8 LoRA(冻结其余权重)可显著扩展链式推理能力,如Qwen3-8B准确率从15.5%升至99%,并提升MuSiQue表现。
AI 中文摘要
预训练Transformer在遵循上下文引用时仅使用了其深度的一小部分。13个基础模型只能可靠地遵循1.4-3.6行,额外的预训练循环带来的提升甚微。在某一早期层上添加一个任务训练的秩为8的LoRA,并在所有模型权重冻结的情况下,可扩展此计算能力。Qwen3-8B在24行链上的精确准确率从15.5%提升至99%;一个训练时间更长的LoRA可处理50行。Ouro-1.4B在四个循环后达到60行,八个循环后至少达到160行。该LoRA启动了一个中继过程:程序行通过中间层的短范围传递其链标识。冻结的注意力头逐步读取链上更远的位置,而移除父行注意力则会中断该中继。一种冻结模型的测量方法在四个保留模型中的三个上,在容差范围内定位了最后一个有效干预层。任务特定的LoRA也提升了MuSiQue的表现。因此,默认答案低估了通过微小编辑可获得的计算能力。代码和交互式演示可在以下网址获取:https://this https URL
英文摘要
Pretrained transformers use little of their depth to follow references in context. Thirteen base models reliably follow only 1.4-3.6 lines, and extra pretrained loops add little. A task-trained rank-8 LoRA at one early layer extends this computation with all model weights frozen. Qwen3-8B improves from 15.5% to 99% exact accuracy on 24-line chains; a longer-trained LoRA reaches 50 lines. Ouro-1.4B reaches 60 lines after four loops and at least 160 after eight. The LoRA starts a relay: program lines pass on their chain identity through a short range of middle layers. Frozen heads read progressively further up the chain, and removing parent-line attention stops the relay. A frozen-model measurement locates the last useful intervention layer within tolerance in three of four held-out models. Task-specific LoRAs also improve MuSiQue. Default answers therefore understate the computation accessible through a tiny edit. Code and an interactive demo are available at https://lunamos.github.io/stop-thinking-too-early/