arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Arkios:从零训练的开放英尼双语语言模型,带有天城文感知分词器

Arkios: An Open Bilingual English-Nepali Language Model Trained From Scratch, with a Devanagari-Aware Tokenizer

Sajal Regmi, Siddhartha Pudasaini, Chetan Phakami Pun

arXiv 2608.30092首次发表:更新:

发表机构

Karela Technologies Inc.(卡雷拉科技公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出从零训练的10.4亿参数英尼双语语言模型Arkios,其在ARC数据集上表现优于同规模模型,揭示了低资源语言模型评估的格式偏差问题,并开源了模型权重。

AI 中文摘要

我们提出Arkios,这是一个10.4亿参数的密集Transformer模型,使用自定义单文件C/CUDA训练栈和为该项目构建的天城文感知字节级BPE分词器,在1500亿个英尼双语文本标记上从零开始预训练。在ARC-Easy和ARC-Challenge数据集上,尽管训练标记数量少了一个数量级,Arkios仍超过了三个规模相当的开放模型(Pythia-1.4B、TinyLlama-1.1B、OLMo-1B),这可能得益于我们的教育网页文本预训练数据与ARC的中小学科学格式相匹配,而非通用能力优势。我们报告了标准协议下的完整评估结果,包括对早期部分样本估计的修正,以及针对低资源语言中小模型评估的特定发现:常用评估工具使用的标准多选字母提示格式使该模型在尼泊尔语阅读理解上处于随机水平,同时在英语上也处于相同格式的随机水平,这会导致一次简单的基准测试运行得出模型没有尼泊尔语能力的结论,而实际上它具备该能力。具体而言,两种语言在字母选择格式下均得分为随机水平(尼泊尔语0.240,英语0.236,随机基线为0.250),而直接对答案文本评分则显示出真实的、偏向英语的理解能力(尼泊尔语0.306,英语0.387)。我们描述了在指令调优期间引入的清单条件工具使用契约,其中仅当上下文声明工具清单时才允许工具调用,否则禁止,并报告了该契约生效和不生效的情况。我们在Apache-2.0许可下发布了基础模型和指令调优后的模型权重;训练代码和尼泊尔语预训练语料库的一小部分私人来源内容未发布;从发布的权重中复现报告数字所需的所有内容均包含在此处。

英文摘要

We present Arkios, a 1.04B-parameter dense transformer pretrained from scratch on 150B tokens of bilingual English-Nepali text, using a custom single-file C/CUDA training stack and a Devanagari-aware byte-level BPE tokenizer built for this project. On ARC-Easy and ARC-Challenge, Arkios exceeds three comparably sized open models (Pythia-1.4B, TinyLlama-1.1B, OLMo-1B) despite an order of magnitude fewer training tokens, likely aided by a match between our educational-web-text pretraining data and ARC's grade-school-science format rather than a general capability advantage. We report full evaluation results under standard protocols, including a correction to an earlier partial-sample estimate, and findings specific to evaluating small models in a low-resource language: the standard multiple-choice-letter prompt format used by common evaluation harnesses places this model at chance on Nepali reading comprehension, and simultaneously at chance on English in the same format, which would lead a naive benchmark run to conclude the model has no Nepali ability when in fact it does. Concretely, both languages score at chance in the letter-choice format (0.240 Nepali, 0.236 English, against a chance baseline of 0.250), while scoring the answer text directly reveals genuine, English-favoring comprehension (0.306 Nepali, 0.387 English). We describe a manifest-conditioned tool-use contract introduced during instruction tuning, where tool calls are permitted only when a tool manifest is declared in context and suppressed otherwise, and report where that contract holds and where it does not. We release both the base and instruction-tuned model weights under Apache-2.0. The training code and a small privately-sourced portion of the Nepali pretraining corpus are not released; everything needed to reproduce the reported numbers from the released weights is included here.

Comments7 pages, 6 tables. Companion paper (tokenizer): arXiv:2608.26449

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑