MaRK:状态空间模型中用于动态算子条件化的马尔可夫自适应循环核
MaRK: Markov-adapted Recurrent Kernels for Dynamic Operator Conditioning in State Space Models
浏览论文内容
中文总结 AI 辅助
MaRK通过将上下文映射为冻结SSM参数的有界调制,实现动态算子条件化,以低参数量(6.3-11M)将双向模型转为迭代扩散,切比雪夫变体最佳(损失2.55)。
中文摘要 AI 辅助
状态空间模型(SSM)为序列建模提供了一种比Transformer更高效的替代方案,然而,对预训练的SSM进行条件化以用于迭代生成,通常是在循环算子之外进行的,通过输入注入或激活调制来实现。虽然这类机制使模型能够接触到条件信息,但底层的时间动态仍然保持不变。我们引入了MaRK(马尔可夫自适应循环核),这是一种动态算子条件化框架,它将上下文向量直接映射为对冻结的SSM的循环($A$)、读入($B$)、读出($C$)、跳跃($D$)和离散化($\Delta$)参数的有界调制。通过LPV-SSM系统的视角,MaRK引出了一个上下文索引的马尔可夫参数序列族,使得每个扩散时间步都能重塑模型的输入-输出记忆核。我们在一个冻结的111M参数Hydra SSM骨干网络上实例化了MaRK,并研究了三种适配器几何结构:Hypernet、切比雪夫多项式和离散余弦变换核。由于这些适配器通过冻结骨干上的低秩辅助映射来修改马尔可夫参数序列,参数高效微调便作为适配机制本身的结构性结果而出现,仅需6.3至11M个可训练辅助参数即可从双向目标过渡到迭代扩散机制。有界的循环参数化进一步为调制循环提供了解析的仿射二次稳定性证书。通过合成LPV恢复实验和马尔可夫算子诊断,我们表明MaRK在匹配假设下能恢复坐标不变的时间算子,并产生不同且稳定的时间步条件化记忆轮廓。实验上,切比雪夫变体表现最强,平均验证损失为2.55,其次是DCT(2.59)和Hypernet(3.77)几何结构。
英文摘要
State Space Models (SSMs) offer an efficient alternative to Transformers for sequence modeling, yet conditioning pre-trained SSMs for iterative generation typically operates outside the recurrent operator, through input injection or activation modulation. While such mechanisms expose the model to conditioning information, they leave the underlying temporal dynamics fixed. We introduce MaRK (Markov-adapted Recurrent Kernels), a dynamic operator-conditioning framework that maps context vectors directly into bounded modulations of a frozen SSM's recurrence ($A$), read-in ($B$), read-out ($C$), skip ($D$), and discretization ($Δ$) parameters. Viewed through the lens of LPV-SSM systems, MaRK induces a context-indexed family of Markov parameter sequences, allowing each diffusion timestep to reshape the model's input-output memory kernel. We instantiate MaRK on a frozen 111M-parameter Hydra SSM backbone and study three adapter geometries: Hypernet, Chebyshev polynomial, and Discrete Cosine Transform kernels. Since these adapters modify the Markov parameter sequence through low-rank auxiliary maps on the frozen backbone, parameter-efficient fine-tuning arises as a structural consequence of the adaptation mechanism itself, requiring only 6.3--11M trainable auxiliary parameters to transition from a bidirectional objective to an iterative diffusion regime. The bounded recurrence parameterization further yields an analytic Affine Quadratic Stability certificate for the modulated recurrence. Through synthetic LPV recovery experiments and Markov-operator diagnostics, we show that MaRK recovers coordinate-invariant temporal operators under matched assumptions and produces distinct, stable timestep-conditioned memory profiles. Empirically, the Chebyshev variant yields the strongest performance, achieving an average validation loss of 2.55, followed by the DCT (2.59) and Hypernet (3.77) geometries.
发表机构
- City University of Hong Kong(香港城市大学)
- NVIDIA(英伟达)
机构由 AI 辅助整理,请以论文原文为准。