发表机构
Institute of Science Tokyo(东京科学研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文扩展Mamba-3,引入输入依赖的低秩反射项以支持非交换状态跟踪,在离散和连续任务上优于标准Mamba-3,实现高精度跟踪。
AI 中文摘要
从序列观测中进行状态跟踪可能既需要保留信息,又需要通过组合观测到的操作来更新信息。我们将Mamba-3的对角转移扩展为输入依赖的低秩反射项,以支持非交换状态跟踪,其中操作的顺序至关重要。秩一更新沿输入依赖方向耦合状态坐标,使得在单个Mamba-3块内能够进行非对角状态转移。该扩展保留了Mamba-3的指数梯形离散化、旋转嵌入(RoPE)和读出机制。在训练方面,我们调整了分块计算以在每个块内并行化所提出的递推关系。实验涵盖了具有离散输入的组词问题和具有连续观测的壳牌游戏,其中策略通过行为克隆进行训练。在固定时间下选出的表现优异的模型中,所提出的模型在具有连续观测和时间抖动的壳牌游戏中,在更长的交换序列上保持了更高的跟踪成功率。这些实验表明,所提出的方法在评估的非交换跟踪任务上实现了高精度,优于标准的Mamba-3。因此,该扩展为基于Mamba-3的非交换状态跟踪提供了一种方法。
英文摘要
State tracking from sequential observations can require both retaining information and updating it by composing observed operations. We extend Mamba-3's diagonal transition with an input-dependent low-rank reflection term to support noncommutative state tracking, in which the order of operations matters. The rank-one update couples state coordinates along an input-dependent direction, enabling non-diagonal state transitions within a single Mamba-3 block. The extension preserves Mamba-3's exponential-trapezoidal discretization, rotary embeddings (RoPE), and readout. For training, we adapt chunkwise computation to parallelize the proposed recurrence within each chunk. Experiments cover group word problems with discrete inputs and a shell game with continuous observations, in which a policy is trained by behavioral cloning. Among the models selected for their strong performance under fixed timing, the proposed model maintains higher tracking success on longer swap sequences in the shell game with continuous observations and timing jitter. These experiments show that the proposed method achieves high accuracy on the evaluated non-commutative tracking tasks, improving on standard Mamba-3. The extension thus offers a Mamba-3-based approach to non-commutative state tracking.
Comments12 pages, 4 figures. Updated experimental results and discussion. Added IEEE submission notice