arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

选择性状态空间模型中状态使用的精确工具及其揭示的输入驱动迁移

An Exact Instrument for State Usage in Selective State-Space Models, and the Input-Driven Migration It Reveals

Raktim Bhattacharya

arXiv 2607.11796首次发表:更新:

发表机构

Texas A&M University(德克萨斯农工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究选择性状态空间模型中状态使用,给出精确测量工具,能计算丢弃模式子集的输出误差。通过该工具发现模型依输入重分配状态空间,输入调度模式剪枝在各规模表现优,揭示了模型状态使用规律及有效剪枝方法。

AI 中文摘要

选择性状态空间模型,如曼巴模型,通过一组一阶模式路由信息,其输入耦合由学习到的选择机制设置。我们给出了一个精确工具来测量训练模型如何使用这些模式。由于状态矩阵是对角的,每个通道的输出可精确分解为每个模式的贡献,并且一个(层,通道,窗口)的Gram张量可离线计算在任何预算下丢弃任何模式子集的精确输出误差。该工具在曼巴 - 1系列上与参考实现验证,相对误差为$2.3×10^{-7}$,在4464种配置上预测层的部署剪枝误差的中值相对偏差为$5×10^{-7}$。在多个模型上应用该工具发现,训练模型会根据输入重新分配其状态空间,输入调度模式剪枝在各规模上优于其他方法。

英文摘要

Selective state-space models such as Mamba route information through a bank of first-order modes whose input coupling is set by a learned selection mechanism. We give an exact instrument for measuring how a trained model uses these modes. Because the state matrix is diagonal, each channel's output decomposes exactly into per-mode contributions, and a per-(layer, channel, window) Gram tensor yields the exact output error of dropping any subset of modes, offline, at any budget. Validated against the reference implementation to a relative error of $2.3\times10^{-7}$ on the Mamba-1 family where it is exact, the instrument predicts a layer's deployed pruning error to a median relative deviation of $5\times10^{-7}$ over $4{,}464$ configurations, its floor set by the reconstruction. Applying the instrument across the Mamba-1 family (130M--2.8B), the deployed 7B Falcon-Mamba, and Mamba-2, we find that trained models re-allocate their state space with the input: which modes carry the signal migrates across contexts, and at the most affected layers a per-input oracle roughly halves the output error of a fixed mode set. Frozen-signal counterfactuals attribute the migration primarily to the input-dependent write map $B_t$; the timestep usually identified with selectivity carries almost none of it. Input-scheduled mode pruning on this measurement outperforms static, Hankel-based, and layer-adaptive rankings at every scale from 130M to the deployed 7B Falcon-Mamba, and at half the state budget it matches the unpruned model. Because the scheduler reads each window's mode usage from a first pass, this demonstrates realizable headroom; we claim no deployed compute or memory saving.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑