AI 中文总结
本研究提出专为CPU推理设计的Daedalus-150M混合架构,在5项任务基准得分超预设线,击败多款更大数据训练的模型,解码速度随上下文长度增长优势显著。
AI 中文摘要
小型语言模型通常仿照大型模型构建,之后再压缩到CPU上运行,而本研究采用相反思路:先确定目标——单用户、逐token处理、4比特权重、普通CPU,再选择适配的架构。最终模型的18个模块中仅6个保留全注意力,其余12个采用短卷积,无论对话多长,其内存仅为2个时间步宽,因此网络的三分之二部分无需重新读取不断增长的缓存。该模型在599亿token数据上从头训练,在5项任务基准测试中得分47.31,而训练前设定的基准线为42.20;它击败了在3至6倍更多数据上训练的GPT-2 124M、Pythia-160M、OPT-125M和GPT-neo-125M,尽管MobileLLM-125M接触过万亿token数据,本模型仍超过其公布的得分。验证比特每字节为0.8685。为验证是架构而非训练方案的作用,本研究在相同数据上训练了同等规模的传统全注意力模型,并在评分前确定了获胜条件。该混合架构在选定质量指标上领先0.81%,在下游任务上表现相当,生成的4比特文件小6.3%,在2048个token上下文下解码速度快1.76倍,与同等规模的外部模型相比快2.08倍。在所有测量中,速度优势在空上下文时接近零,随上下文长度增加而增大,这符合机制预测,单纯精简模型不会出现此情况。简单带宽计算仅预测1.17倍,仅内存容量无法解释差距。本研究还报告了未成功的尝试:纯粹的4比特质量损失、约一半的卷积通道失效且无法移除、词汇量超出该模型规模应有的大小。
英文摘要
Small language models are usually built like large ones and then squeezed onto a CPU afterwards. We did the opposite: we fixed the target first, one user, one token at a time, 4-bit weights, ordinary CPU, and chose the architecture to suit it. The result keeps full attention in only 6 of its 18 blocks. The other 12 use short convolutions whose memory is two timesteps wide no matter how long the conversation gets, so two thirds of the network never re-reads a growing cache. Trained from scratch on 59.9B tokens, the model scores 47.31 on a five-task benchmark against a bar of 42.20 that was fixed before training began. It beats GPT-2 124M, Pythia-160M, OPT-125M and GPT-neo-125M, all trained on three to six times more data, and exceeds MobileLLM-125M's published score despite that model seeing a trillion tokens. Validation bits-per-byte is 0.8685. To check the architecture rather than the training recipe, we trained a conventional all-attention model of the same size on the same data, and wrote down the winning condition before scoring either. The hybrid won the chosen quality metric by 0.81%, matched it on downstream tasks, produced a 6.3% smaller 4-bit file, and decoded 1.76x faster at 2048 tokens of context, 2.08x against an external model of similar size. In every measurement the speed advantage is near zero at an empty context and grows with length, which is what the mechanism predicts and what a merely leaner model would not show. A simple bandwidth calculation predicts only 1.17x, so memory volume alone does not explain the gap. We also report what did not work: an unmitigated 4-bit quality cost, roughly half the convolution channels ending up inert and impossible to remove, and a vocabulary larger than this model size warrants.
Comments8 pages, 2 figures, 10 tables. Code, weights and the measurement artefacts behind every number: https://github.com/unseen1980/daedalus