发表机构
Oklahoma Christian University(俄克拉荷马基督大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对小参数语言模型中输出矩阵占用容量大的问题,提出基于黎曼流形测地线距离解码的RiLM框架,移除输出层,在WikiText-2上以29万参数实现54.2的困惑度,较最强基线提升约2倍。
AI 中文摘要
参数低于一百万个的语言模型对于边缘部署、领域适应和可复现研究至关重要,然而一个两层LSTM或Transformer在嵌入宽度d=128时,仍会将其约三分之一的容量花费在输出矩阵W_out(维度为R^(d x |V|))上。我们提出了黎曼语言模型(RiLM),它完全移除了该层:上下文在黎曼流形上展开为一条轨迹,下一个词元的概率由当前状态与词汇嵌入之间的平方测地线距离得出。同一个嵌入映射同时用于输入和输出——解码即几何。我们在平坦空间R^d(Flat RiLM)和庞加莱球H^d(HypRiLM)上实例化了该框架,并采用共享的MLP组合映射phi(约29万参数,d=128,|V|=2000)。在WikiText-2上跨越五个随机种子,HypRiLM达到54.2±0.2的验证困惑度,而Flat RiLM为87.6±0.6;绑定的匹配LSTM、Transformer和SSM基线保持在113-147的困惑度(WT-2上)——HypRiLM比最强的绑定循环基线(SSM,113.0±3.8)领先约2倍。Penn Treebank和一个10k词汇量的压力测试证实,测地线解码可跨语料库和更大的|V|迁移,而双曲曲率具有选择性帮助。我们还刻画了朴素双曲循环中的边界坍缩问题,并展示了莫比乌斯稳定化如何恢复可训练性。我们的主张限定于受控的小模型比较,而非全词汇量的最先进水平。
英文摘要
Language models under one million parameters matter for edge deployment, domain adaptation, and reproducible research, yet a two-layer LSTM or Transformer at embedding width d = 128 still spends roughly one third of its capacity on the output matrix W_out in R^(d x |V|). We propose Riemannian Language Models (RiLM), which remove that layer entirely: context unfolds as a trajectory on a Riemannian manifold, and next-token probabilities arise from squared geodesic distance between the current state and vocabulary embeddings. The same embedding map serves input and output -- decoding is geometry. We instantiate the framework on flat R^d (Flat RiLM) and the Poincare ball H^d (HypRiLM) with a shared MLP composition map phi (~290k parameters, d = 128, |V| = 2000). Across five seeds on WikiText-2, HypRiLM reaches 54.2 +/- 0.2 validation perplexity versus 87.6 +/- 0.6 for Flat RiLM; tied and matched LSTM, Transformer, and SSM controls remain at 113-147 PPL on WT-2 -- HypRiLM leads by roughly 2x over the strongest tied recurrent baseline (SSM, 113.0 +/- 3.8). Penn Treebank and a 10k-vocabulary stress test confirm that geodesic decoding transfers across corpora and larger |V|, while hyperbolic curvature helps selectively. We also characterize boundary collapse in naive hyperbolic recurrence and show how Mobius stabilization restores trainability. Claims are scoped to controlled small-model comparisons, not full-vocabulary state of the art.