写回 $\Delta$:以全新表示重新审视相同 Token
Write Back the $Δ$: Revisiting the Same Tokens with Fresh Representations
浏览论文内容
中文总结 AI 辅助
提出 ReFlux,一种可学习的反馈图,通过将层间增量 $\Delta$ 写回早期层,使 Transformer 以全新表示重新审视相同 Token,在多个基准上降低困惑度并提升准确率,同时保持计算开销。
中文摘要 AI 辅助
Transformer 在深度方向上严格地前向处理信息,这阻止了更深层的计算重新审视并细化早期表示。为了增强标准的前向传播,现有方法要么重新执行深度计算(导致额外计算开销),要么使用预定义方向修改残差流(限制了其实例级适应性)。最近,推理时反馈提供了一种直接机制,通过将更深的残差状态写回较早层来循环利用内部产生的计算,但究竟应该反馈什么仍不清楚。我们认为,深度增量 $\Delta$(捕获两层之间新累积的计算)比完整状态提供了更有效、更可组合且更具可扩展性的反馈信号。基于这一观察,我们引入了 ReFlux,一个可学习的反馈图,它动态选择并组合携带增量的路径。ReFlux 支持对相同 Token 的同步反馈和对后续 Token 的流式反馈。跨多种模型、语料库和基准的大量实验表明,同步 ReFlux 在十个语言建模语料库上持续降低困惑度,并将准确率提高 2.1-2.3 个百分点,在多跳推理上增益达到 4.7 个百分点。流式 ReFlux 进一步保留了大部分这些增益,同时保持基础模型 1 倍的理论主干 FLOPs。这些结果确立了 ReFlux 作为一种高效范式,用于释放 LLM 的潜在计算能力,使它们能够以全新表示重新审视相同 Token。代码实现可在 https://this https URL 找到。
英文摘要
Transformers process information strictly forward through depth, preventing deeper computation from revisiting and refining earlier representations. To augment the standard forward pass, existing approaches either re-execute depth, incurring additional computation, or modify the residual stream using predefined directions, limiting their instance-level adaptation. Recently, inference-time feedback offers a direct mechanism for recycling endogenously produced computation by writing deeper residual states back to earlier layers, yet what should be fed back remains unclear. We argue that the depth increment Delta, capturing newly accumulated computation between two layers, provides a more effective, composable, and scalable feedback signal than the full state. Building on this observation, we introduce ReFlux, a learnable feedback graph that dynamically selects and composes increment-carrying routes. ReFlux supports synchronous feedback to the same token and streaming feedback to subsequent tokens. Extensive experiments across various models, corpora, and benchmarks show that synchronous ReFlux consistently reduces perplexity across ten language-modeling corpora, and improves accuracy by 2.1-2.3 points, with gains reaching 4.7 points on multi-hop reasoning. Streaming ReFlux further retains most of these gains while preserving the base model's 1x theoretical backbone FLOPs. These results establish ReFlux as an efficient paradigm for unlocking the latent computational potential of LLMs, allowing them to revisit the same tokens with fresh representations. Code implementation can be found at https://github.com/gooogleshanghai/reflux.