发表机构
University of Bonn; Lamarr Institute(波恩大学; 拉马尔研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过拆分前向传播为写入器和读取器,发现语言模型在操作token的KV缓存中留下局部可恢复的记录,支持路由和直接载荷访问,但模型主要读取地址而非值。
AI 中文摘要
当语言模型读取诸如“交换盒子F和盒子B的内容”这样的操作时,其前向传播会为这些token将键和值写入KV缓存。先前关于实体追踪的研究确立了模型使用的内容:绑定在查询时被解析,而非存储为显式的潜在状态。我们探究它们在操作跨度内写入了什么,以及这些内容是如何被访问的。我们将一次前向传播拆分为一个冻结的写入器和一个读取器:写入器的缓存无需梯度即可重新计算,而读取器仅能看到指令和操作token,所有状态描述均被隐藏,并单独进行训练。因此,读取器恢复出的任何内容必然已存在于未修改的缓存中。在一个合成的盒子任务上,基础读取器在训练前仅能恢复出所查询绑定的$\leq 0.06$,而训练后恢复率可达$0.75$--$1.00$,且可恢复性与操作的读/写足迹相关。我们发现了两种访问模式。在Llama-3.1-8B和Mistral-7B上,操作跨度移植能因果性地重定向所读取的可见状态,即使两个世界持有相同的值,这揭示了一种路由记录。隔离训练保留了路由功能,并增加了对载荷(即操作读取的值)的直接访问,该访问来自一个狭窄的中深度层带(Llama-3.1-8B的32层中的第12--15层,Mistral-7B中的第14--17层)中的单个操作数名称token——该位置正是持有路由记录之处。同样的方法可扩展到更多操作、ToMi和GSM8K,但受限于训练覆盖范围,并会牺牲开放书准确性。因此,操作token会留下局部的、可因果恢复的记录,这些记录既支持路由也支持直接载荷访问,尽管写入这些记录的模型主要读取的是它们携带的地址,而非值本身。
英文摘要
When a language model reads an operation such as "Swap the contents of Box F and Box B", its forward pass writes keys and values for those tokens into the KV cache. Prior work on entity tracking establishes what models use: bindings are resolved at query time rather than stored as explicit latent state. We ask what they write at the operation span and how it is accessed. We split a forward pass into a frozen writer and a reader: the writer's cache is recomputed without gradients, while the reader sees only the instruction and operation tokens, with all state descriptions hidden, and is trained in isolation. Anything the reader recovers was therefore already present in the unmodified cache. On a synthetic boxes task, a base reader recovers $\leq 0.06$ of queried bindings against $0.75$--$1.00$ after training, and recoverability tracks the operation's read/write footprint. We find two modes of access. Across Llama-3.1-8B and Mistral-7B, operation-span transplants causally redirect which visible state is read even when the two worlds hold identical values, revealing a routing record. Isolation training preserves routing and adds direct access to the payload, the value the operation read, from the single operand-name token in a narrow mid-depth band (layers 12--15 of 32 in Llama-3.1-8B, 14--17 in Mistral-7B) --- the same site that holds the routing record. The same recipe extends to further operations, ToMi and GSM8K, but is bounded by training coverage and costs open-book accuracy. Operation tokens thus leave localized, causally recoverable records that support both routing and direct payload access, though the model that writes them reads mainly the address they carry and not the value.