发表机构
SAP(SAP)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
对TALH混合语言模型进行消融实验,发现移除SSM分支影响最大,而移除MLA影响较小,结果支持特定实现假设而非通用结论。
AI 中文摘要
我们报告了一项探索性的、单一种子(single-seed)的消融研究,研究对象为TALH(自适应潜在混合,Adaptive Latent Hybrid),一种仅解码器(decoder-only)的语言模型,其具有并行多头潜在注意力(Multi-head Latent Attention, MLA)和自定义循环状态空间(state-space, SSM)分支。我们训练了五个变体,每个token的估计活跃参数在1.17亿至2.17亿之间,在FineWeb样本上从头开始训练,使用相同的优化步数和token数量。在这种特定设置下,移除SSM分支会导致验证困惑度(perplexity)的最大下降(仅MLA的困惑度为315),而移除MLA的影响则小得多(仅SSM的困惑度为239)。一个密集前馈网络(dense-FFN)混合模型获得了231的困惑度,而测试的top-2三元混合专家(ternary-MoE)混合模型的困惑度为240,同时峰值训练内存减少了3.87 GB。我们还保留了一个初步的Apple M3计时观察结果:在五个未优化的实现中,仅MLA模型在512到2,048个提示token范围内的首token生成时间(time-to-first-token)曲线最为平缓,尽管密集Transformer在绝对时间上要快得多。由于这些运行是单一种子的,参数数量不匹配,评估流可能与训练源重叠,且缺少原始重复计时记录,因此这些结果支持特定于实现(implementation-specific)的假设,而非关于MLA、SSM或混合专家模型的通用结论。
英文摘要
We report an exploratory, single-seed ablation of TALH (Adaptive Latent Hybrid), a decoder-only language model with parallel Multi-head Latent Attention (MLA) and a custom recurrent state-space (SSM) branch. Five variants, spanning 117--217M estimated active parameters per token, are trained from scratch on a FineWeb sample for the same number of optimisation steps and tokens. In this specific setup, removing the SSM branch gives the largest degradation in validation perplexity (MLA-only PPL 315), whereas removing MLA has a much smaller effect (SSM-only PPL 239). A dense-FFN hybrid obtains PPL 231, compared with 240 for the tested top-2 ternary-MoE hybrid, while using 3.87 GB less peak training memory. We also preserve a preliminary Apple M3 timing observation: among the five unoptimised implementations, MLA-only has the flattest measured time-to-first-token curve from 512 to 2,048 prompt tokens, although the dense Transformer is much faster in absolute terms. Because the runs are single-seed, parameter counts are unmatched, the evaluation stream may overlap the training source, and raw repeated timing records are unavailable, these results support implementation-specific hypotheses rather than general conclusions about MLA, SSMs, or mixture-of-experts models.
Comments6 pages, 7 figures. Exploratory single-seed ablation study. Code and replication package available at https://github.com/unseen1980/talh