ufakzeka-1:从零构建并评估一个151M参数的土耳其语语言模型
ufakzeka-1: Building and Evaluating a 151M-Parameter Turkish Language Model from Scratch
浏览论文内容
中文总结 AI 辅助
本文从零构建并评估151M参数的土耳其语模型ufakzeka-1,提出分词器、三阶段预训练及评估套件,发现安全门修复假象、种子方差大及模型规模限制等问题。
中文摘要 AI 辅助
我们描述了ufakzeka-1,一个151M参数(含嵌入层为182M)的仅解码器土耳其语语言模型,该模型从零开始在135亿个开放许可文本的token上进行了预训练,并针对聊天进行了指令微调,总成本约为286美元的云GPU、API和笔记本时间。其贡献不在于模型的能力——这是该规模模型应有的表现——而在于构建和评估它的记录:一个土耳其语字节级分词器,每词1.77个token;一个三阶段预训练计划;一个由开放许可和生成数据组成的后训练混合;以及一个由发布门控、对5,508个对话进行规则检查的扫描、人工评判的对话和手工测试组成的评估套件,所有提示均从训练数据中保留,通过数据构建中的去污染和我们在每次构建前运行的已检查不变式脚本强制执行。我们报告了我们认为可迁移到其他小模型工作的三个发现:一个曾被其自身问题生成的训练数据“修复”的安全门读数为64/64,而真实数字为34/64;训练种子的方差与我们尝试的每种配方的差异一样大,因此在此规模下单一种子比较无信息量;数据轮次仅修复了数据中缺失的内容,而长上下文中的身份追踪和多轮算术在我们尝试的任何数据变更中均未移动,我们将其解读为模型规模的限制而非数据缺口,这一解读将由下一个更大的模型来检验。权重、数据配方、评估代码和支出账本均在Apache-2.0许可下发布。
英文摘要
We describe ufakzeka-1, a 151M-parameter (182M with embeddings) decoder-only Turkish language model pretrained from scratch on 13.5B tokens of openly licensed text and instruction-tuned for chat, at a total cost of about \$286 in cloud GPU, API and notebook time. The contribution is not the model's capability, which is what a model this size can be expected to have, but the record of building and measuring it: a Turkish byte-level tokenizer at 1.77 tokens per word, a three-stage pretraining schedule, a post-training mixture of openly licensed and generated data, and an evaluation battery of release gates, a rule-checked sweep of 5,508 conversations, judged conversations and hand tests, all with prompts held out from the training data, enforced by decontamination inside the data build and by a checked-in invariant script we run before each build. We report three findings that we believe transfer to other small-model efforts: a safety gate that had been "fixed" with training data written from its own questions read 64/64 while the honest figure was 34/64; training-seed variance was as large as the spread across every recipe we tried, so single-seed comparisons at this scale are uninformative; and data rounds repaired only what was absent from the data, while identity tracking over long context and multi-turn arithmetic did not move across any data change we tried, which we read as limits of the model size rather than gaps in the data, a reading the next, larger model will test. Weights, the data recipe, the evaluation code and the spend ledger are released under Apache-2.0.
发表机构
- ufak AI
机构由 AI 辅助整理,请以论文原文为准。