残差流读取,循环状态记忆:Mamba模型中的全局工作空间
Residual Streams Read, Recurrent States Remember: The Global Workspace in Mamba Models
浏览论文内容
中文总结 AI 辅助
本研究探讨全局工作空间理论在Mamba状态空间模型中的适用性,通过Jacobian透镜拟合残差流与循环状态,发现联合读出提升概念恢复,循环状态可独立支持口头访问,并提出符号保护引导方法,但关系性答案增益有限。
中文摘要 AI 辅助
变压器表示中的全局工作空间理论能否扩展到状态空间语言模型?我们使用原始的1000提示配方,将Jacobian透镜拟合到Mamba-1、Mamba-2和Mamba-3的残差流和循环状态上。在每一个测试的Mamba检查点上,联合残差-状态读出在至少六个任务家族中的五个上,比残差透镜更好地恢复已知的中间概念。在Mamba-2上,仅状态在所有六个家族上超过了残差和logit透镜;归一化联合读出在五个家族上优于两个组成部分。时间映射和词表实验表明,随着残差可见性的变化,早期内容仍可被状态读取。我们还提出了符号保护的引导(sign-guarded steering),在五个模型的匹配口头报告试验中,相比坐标交换,提高了目标前五成功率。仅循环状态就支持这种口头访问。这些增益在关系性答案上并不一致地扩展:受保护的编辑常常输出编辑后的概念本身,而Mamba-3的联合编辑可能破坏仅状态重定向的成功。因此,循环状态提供了工作空间内容的互补载体,其恢复、持久性和因果用途需要单独测量。
英文摘要
Can the global-workspace account of transformer representations extend to state-space language models? We fit Jacobian lenses to the residual streams and recurrent states of Mamba-1, Mamba-2 and Mamba-3, using the original 1000-prompt recipe. Joint residual--state readouts improve recovery of known intermediate concepts over the residual lens on at least five of six task families in every tested Mamba checkpoint. On Mamba-2, state alone exceeds residual and logit lenses on all six families; a normalised joint readout improves on both components on five. Temporal maps and word-list experiments show earlier content remaining state-readable as residual visibility changes. We also propose sign-guarded steering, which improves target top-five success over coordinate exchange on matched verbal-report trials in five models. Recurrent state alone supports this verbal access. These gains do not extend consistently to relational answers: guarded edits often output the edited concept itself, and Mamba-3's joint edits can disrupt successful state-only redirection. Recurrent state thus provides a complementary carrier of workspace content, whose recovery, persistence and causal uses require separate measurements.