发表机构
The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出SCIT因果协议,测试潜在思维链模型的缓存载体,发现反事实算术主要通过值缓存后缀轨迹传递,且载体机制随模型规模和类型存在差异。
AI 中文摘要
潜在思维链模型将中间推理从生成文本转移到连续状态,提升了紧凑性但隐藏了因果对象。我们提出SCIT,即后缀缓存互换测试,这是一种因果协议,可构建精确的源-接收者反事实、修补声明的缓存片段,并识别哪个Transformer对象承载反事实计算。SCIT结合了充分性测试与键/值(K/V)组件拆分、隐藏状态控制、语义源控制、解码验证及匹配损坏。在CODI-GPT2和Sim-CoT风格的GPT-2复现模型上,反事实算术主要通过值缓存后缀轨迹传递,而非隐藏状态、键、可重用答案槽或单令牌触发器。对于主要的CODI-GPT2检查点,存在值后缀机制的充分性和必要性完整证据;Sim-CoT风格检查点显示相同的充分性和解码控制模式,但缺乏匹配损坏的必要性证据。超出这些局部算术单元,SCIT揭示了载体机制转变:类算术的GPT-2/1B单元保留潜在尾部值/KV传递,而具备能力的8B单元及修复后的非算术单元通过提示前缀或完整缓存KV传递;边界单元未接收机制调用。因此,SCIT提供了缓存级诊断、特定检查点的GPT-2算术机制,以及能力门控的载体映射,而非通用潜在尾部主张。
英文摘要
Latent chain-of-thought models move intermediate reasoning from emitted text into continuous states, improving compactness but hiding the causal object. We introduce SCIT, the Suffix Cache Interchange Test, a causal protocol that constructs exact source-recipient counterfactuals, patches declared cache segments, and identifies which transformer object carries the counterfactual computation. SCIT combines sufficiency tests with K/V component splits, hidden-state controls, semantic source controls, decoded validation, and matched corruption. On CODI-GPT2 and a Sim-CoT-style GPT-2 reproduction, counterfactual arithmetic transfers primarily through value-cache suffix trajectories rather than hidden states, keys, reusable answer slots, or single-token triggers. Complete sufficiency-and-necessity evidence for the late-value-suffix mechanism holds for the main CODI-GPT2 checkpoint; the Sim-CoT-style checkpoint shows the same sufficiency and decoded-control pattern but insufficient matched-corruption evidence for a necessity call. Beyond these local arithmetic cells, SCIT reveals carrier-regime shifts: arithmetic-like GPT-2/1B cells preserve latent-tail value/KV transfer, whereas competent 8B and repaired non-arithmetic cells route through prompt-prefix or full-cache K/V; boundary cells receive no mechanism call. SCIT therefore contributes a cache-level diagnostic, a checkpoint-specific GPT-2 arithmetic mechanism, and a competence-gated carrier map rather than a universal latent-tail claim.
Commentsaccept by emnlp2026