arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32486cs.LG

弹性选择性频谱混合模型:用于一次训练、多次导出的预算推理

Elastic Selective Spectral Hybrids for Train-Once, Export-Many Budgeted Inference

Dachuan Song, Chuchu Chen, Xuan Wang

AI总结:

针对不同预算下语言模型部署,提出弹性选择性频谱混合模型(ESSH),通过选择性频谱混合器与滑动窗口注意力结合,实现一次训练多容量导出,在保持质量的同时提供显著推理加速。

AI中文摘要:

在不同计算和延迟预算下部署语言模型,需要具有不同质量-成本权衡的紧凑模型。为此,弹性频谱状态空间模型提供了有序且时间分解的信道,这些信道可以被截断,但它们与线性时不变滤波器相关联,无法在上下文演变时选择性地保留相关过去信息或遗忘无关信息。为解决这一问题,我们引入了弹性选择性频谱混合模型(ESSH),该模型将每个Hankel频谱信道实现为使用拟合阻尼旋转模式的独立循环单元。它还具有依赖于输入的衰减和写/读门,使时间保留和状态更新依赖于输入,同时保持信道级截断和结构化循环以实现高效执行。ESSH将这些选择性频谱混合器与滑动窗口注意力相结合,并通过双速率容量映射和全模型蒸馏,以不同速率减少频谱信道数量和前馈宽度,从而联合训练多种容量。所得模型支持分块并行训练和融合循环解码,同时避免对丢弃信道的计算。在满容量下,ESSH实现了与类似规模的独立训练模型相当的语言建模质量,而较小的导出模型则表现出平滑的质量-成本权衡。我们使用语言理解、检索、跨域文本和DNA实验验证了所提出框架的有效性,通过评估质量保持以及与独立训练和弹性基线的权衡。在1.53B模型配置下,融合批大小为1的解码在B300上每token耗时1.37毫秒,与测试的Mamba-2和Mamba-3实现相比提供了2.14-2.80倍的加速,在匹配参数数量下比Transformer++快3.03倍。

英文摘要:

Deploying a language model under different computing and latency budgets calls for compact models with different quality-cost trade-offs. Towards this end, elastic spectral state space models provide ordered and temporally decomposed channels that can be truncated, but they are associated with linear time-invariant filters that cannot selectively preserve relevant past information or forget irrelevant information as the context evolves. To address this, we introduce the Elastic Selective Spectral Hybrid (ESSH), which realizes each Hankel spectral channel as an independent recurrent unit using fitted damped rotation modes. It also features an input-dependent decay and write/read gates that make temporal retention and state update input-dependent while preserving channel-wise truncation and a structured recurrence for efficient execution. ESSH combines these selective spectral mixers with sliding-window attention and jointly trains multiple capacities by reducing spectral-channel count and feed-forward width at different rates through a two-rate capacity map with full-model distillation. The resulting models support chunked parallel training and fused recurrent decoding while avoiding computation for discarded channels. At full capacity, ESSH achieves language-modeling quality comparable to similarly sized independently trained models, while smaller exports exhibit a smooth quality-cost trade-off. We validate the effectiveness of the proposed framework using language understanding, retrieval, cross-domain text, and DNA experiments, by assessing quality retention and the trade-off against independently trained and elastic baselines. At the 1.53B model configuration, fused batch-one decoding takes 1.37 ms per token on a B300, providing a 2.14-2.80x speedup over the tested Mamba-2 and Mamba-3 implementations and 3.03x over Transformer++ at matched parameter counts.

↑