掩码并非模型:审计注意力、状态空间与混合序列模型中的前缀不变性
The Mask Is Not the Model: Auditing Prefix Invariance in Attention, State-Space, and Hybrid Sequence Models
浏览论文内容
中文总结 AI 辅助
本研究针对注意力、状态空间及混合序列模型,提出无需训练的轻量级前缀不变性审计方法,可精准定位因果关系破坏点,且在测试中成功发现Zamba2等模型的缺陷。
中文摘要 AI 辅助
我们形式化了前缀不变性:位置t的表示不得依赖未来输入。我们提出一种轻量级审计方法,仅需两次前向传播,无需训练或梯度,即可精准定位因果关系被破坏的位置。注意力掩码检查并不完整:即便掩码正确,仍可能通过扫描或归一化出现信息泄露。在八个检查点的192次注入故障试验中,掩码检查未发现任何问题,而我们的审计方法定位到了全部192/192个故障,还发现了Zamba2和Nemotron-H存在缺陷。
英文摘要
Hybrid sequence models must satisfy prefix invariance: representations at position t must not depend on future inputs, yet this is rarely verified. We formalize prefix invariance and give a lightweight audit, two forward passes, no training or gradients, yielding a per-layer score localizing where causality breaks. Attention-mask inspection, the field's default check, is incomplete: causality is a graph-level property, and leaks can occur via scans, aggregations, or normalization despite correct masks. Across 192 injected-fault trials on eight checkpoints, mask inspection detected none, while our audit localized all 192/192 to the exact layer. Static/dynamic analysis of chunked-scan code in transformers found the same defect in Zamba2 and Nemotron-H, an inter-chunk axis error fixed via the reference implementation. The method fits on one page and runs in seconds.
发表机构
- VIDRAFT AI Research(VIDRAFT AI研究机构)
机构由 AI 辅助整理,请以论文原文为准。