arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

语言模型需要可训练的输入嵌入表吗?1.7B 类规模下的固定最小令牌编码

Do Language Models Need a Trainable Input Embedding Table? Fixed Minimal Token Codes at 1.7B-Class Scale

A. Bochkov

arXiv 2610.04002首次发表:更新:

AI 中文总结

本研究通过对比三种输入接口(可学习表、固定编码)在1.7B规模模型上的表现,证明固定令牌编码无需可训练输入嵌入表即可获得显著语言建模能力,确立了架构可行性。

AI 中文摘要

可训练的输入嵌入表为每个词汇项分配一个独立可调的向量。我们研究这种令牌特定的参数化是否是实现显著语言建模能力所必需的,还是共享的 Transformer 可以从固定的令牌身份中学习。我们比较了三个从头训练的仅解码器语言模型,它们使用相同的分词器、上下文骨干、非共享输出头架构和训练方案,每个模型的目标预算为 1000 亿个预测令牌。它们的输入接口分别是学习到的表、规范的 16 位令牌 ID 编码,以及一种在 GF(2) 上的固定可逆重编码。固定编码被重复到模型宽度,无需额外的可训练输入投影。两个固定编码模型都获得了显著的能力:规范编码在 HellaSwag 上达到 52.40% 的归一化准确率,在 PIQA 上达到 70.51% 的准确率,在 LAMBADA 上达到 42.75% 的准确率。学习输入的控制模型在多项评估中表现更好,包括 HellaSwag 和 LAMBADA,因此这些结果确立了可行性而非性能对等。固定接口移除了 1.007 亿个可训练参数,得到 1.711B 参数的模型,但参数减少并非核心结果。这些单次运行实验区分了架构必要性与经验效用:独立可训练的令牌特定输入向量并非观察到的能力所必需。固定身份接口还为研究不可变输入下游的表示学习提供了一个受控环境,而无需确定特定能力的定位。

英文摘要

A trainable input embedding table assigns each vocabulary item an independently adjustable vector. We investigate whether this token-specific parameterization is required for substantial language-modeling capability, or whether a shared Transformer can learn from fixed token identities. We compare three decoder-only language models trained from scratch with the same tokenizer, contextual backbone, untied output-head architecture, and training recipe, with a target budget of 100 billion prediction tokens per model. Their input interfaces are a learned table, canonical 16-bit token-ID codes, and one fixed invertible recoding over GF(2). The fixed codes are repeated to model width without an additional trainable input projection. Both fixed-code models acquire substantial capabilities: canonical codes achieve 52.40\% HellaSwag normalized accuracy, 70.51\% PIQA accuracy, and 42.75\% LAMBADA accuracy. The learned-input control performs better on several evaluations, including HellaSwag and LAMBADA, so these results establish viability rather than performance parity. The fixed interfaces remove 100.7 million trainable parameters, yielding 1.711B-parameter models, but parameter reduction is not the central result. These single-run experiments distinguish architectural necessity from empirical utility: independently trainable token-specific input vectors are not required for the observed capabilities. A fixed identity interface also provides a controlled setting for studying representation learning downstream of an immutable input, without establishing where particular capabilities are localized.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑