雅可比视角下的循环Transformer:全局工作空间在循环机制下是否依然存在?
Looped Transformers under the Jacobian Lens: Does the Global Workspace Survive Recurrence?
浏览论文内容
中文总结 AI 辅助
该研究以雅可比视角探究循环Transformer的全局工作空间,发现两类迭代架构(Ouro-2.6B、Huginn-0125)的迭代部分均形成工作空间,但循环机制改变了其访问方式,且新内容可言语化与每轮监督相关。
中文摘要 AI 辅助
近期研究在标准前馈Transformer中识别出一段可被言语化、具有因果效力的中层表征带,它是全局工作空间的功能类似物。当通过循环而非堆叠不同层实现深度时,是否会出现相同的工作空间功能尚不清楚。循环与深度循环Transformer可直接检验该问题,因为它们在深度维度复用相同权重。我们使用虚拟展开适配器将雅可比视角扩展至迭代架构,将全套工作空间工具(视角拟合、读出操作及11类因果实验系列)应用于Ouro-2.6B(48层循环4次,深度监督)、Huginn-0125(4层核心循环16次,用于潜在推理训练),并以Qwen3.6-27B(64个未绑定层)作为标准基线。我们发现,每个架构的迭代部分都会形成工作空间,但循环改变了其访问方式:Ouro在每次循环中重构工作空间内容,线性传输无法将内容跨循环边界传递,因此写入与消融操作必须覆盖所有剩余循环;Huginn将内容向前传递至全部16次循环,而读取、写入与消融仅在约2次循环的滑动窗口内生效。新注入内容能否被言语化与显式的每轮监督相关,现有内容能否被调控则与该监督无关。
英文摘要
Recent work identifies a mid-depth band of verbalisable, causally potent representations in a standard feedforward transformer --- a functional analogue of a global workspace. Whether the same workspace functionality emerges when depth is implemented through recurrence rather than a stack of distinct layers remains unknown. Looped and depth-recurrent transformers provide a direct test of this question because they reuse the same weights across depth. We extend the Jacobian lens to iterated architectures using a virtual-unrolling adapter. We apply the full workspace suite --- lens fitting, readout, and eleven causal experiment families --- to Ouro-2.6B (48 layers looped 4 times, deeply supervised) and Huginn-0125 (a 4-layer core recurred 16 times, trained for latent reasoning), using Qwen3.6-27B (64 untied layers) as the standard baseline. We find that a workspace forms in the iterated part of each architecture, but that recurrence changes how it can be accessed. Ouro reconstructs workspace content in every loop, and linear transport cannot carry that content across loop boundaries; writes and ablations must therefore span every remaining loop. Huginn carries content forward across all sixteen recurrences, while reads, writes, and ablations act only within a sliding window of roughly two recurrences. Whether newly injected content can be verbalised tracks explicit per-iteration supervision; whether existing content can be steered does not.
发表机构
- Fin AI Research(Fin AI 研究院)
机构由 AI 辅助整理,请以论文原文为准。