共享相位与保持控制:高效自适应谱递归
Shared Phase and Retention Control for Efficient Adaptive Spectral Recurrence
浏览论文内容
中文总结 AI 辅助
提出SPARC,用两个共享标量控制信号协调谱递归的记忆保持与相位旋转,实现高效自适应记忆,在控制与分类任务上超越现有方法,并显著加速训练。
中文摘要 AI 辅助
随着新证据的到来,序列模型必须更新其记忆内容以及记忆对预测的影响方式。尽管Transformer模型的计算和缓存成本随上下文长度增长,固定状态递归模型则提供恒定内存的推理。然而,线性和谱递归传统上依赖静态转移,无法动态修正存储表示的衰减或旋转方式。虽然近期选择性架构引入了输入依赖的转移,但它们为每个记忆模式分配独立控制,将控制成本与状态容量耦合。我们证明高维谱记忆并不需要高维控制,并引入了共享相位与保持控制的高效自适应谱递归(SPARC)。SPARC仅使用两个输入依赖的标量信号来协调异构复数模式间的记忆保持和相位旋转,同时保留各模式特定的基线时间尺度和频率。其对角仿射递归支持并行关联扫描以实现序列级BPTT,以及精确的结构化实时递归学习(RTRL)用于在线信用分配。在部分可观测连续控制、POPGym和序列分类任务中,SPARC在Walker-P上相对第二优方法实现了9.09%的相对回报提升,在FordA上实现了1.36%的相对准确率提升。在NVIDIA Blackwell GPU上,我们的实现将固定令牌工作负载中的递归混合器训练延迟降低了18.2%-34.2%,并将扫描速度相对于优化的RG-LRU基线加速了3.1倍至4.7倍。这些结果表明,两个共享控制信号能够在在线和全序列设置中高效地管理自适应谱记忆。代码可在https URL获取。
英文摘要
As new evidence arrives, a sequence model must update what it remembers and how memory influences predictions. While Transformers incur computation and cache costs scaling with context length, fixed-state recurrent models offer constant-memory inference. However, linear and spectral recurrences traditionally rely on static transitions, failing to dynamically revise how stored representations decay or rotate. While recent selective architectures introduce input-dependent transitions, they assign independent controls to every memory mode, coupling control cost to state capacity. We show that high-dimensional spectral memory does not require high-dimensional control, and introduce Shared Phase and Retention Control for Efficient Adaptive Spectral Recurrence (SPARC). SPARC employs just two input-dependent scalar signals to coordinate memory retention and phase rotation across heterogeneous complex modes, while preserving mode-specific baseline timescales and frequencies. Its diagonal affine recurrence supports parallel associative scans for sequence-level BPTT as well as exact structured Real-Time Recurrent Learning (RTRL) for online credit assignment. Across partially observable continuous control, POPGym, and sequence classification, SPARC achieves a 9.09% relative return improvement on Walker-P and a 1.36% relative accuracy gain on FordA over second-best methods. On an NVIDIA Blackwell GPU, our implementation reduces recurrent-mixer training latency by 18.2%-34.2% in fixed-token workloads and accelerates scans by 3.1x-4.7x over an optimized RG-LRU baseline. These results show that two shared control signals can efficiently govern adaptive spectral memory across online and full-sequence settings. Code is available at https://github.com/Botwwt/sparc.
发表机构
- Peking University(北京大学)
- GigaAI(极佳科技)
机构由 AI 辅助整理,请以论文原文为准。