arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Manacá-1B:一个开源、可复现的巴西葡萄牙语语言模型及一种感知分词器的配对评估方法

Manacá-1B: An Open, Reproducible Brazilian-Portuguese Language Model and a Tokenizer-Aware, Paired Evaluation

Bruno Leonardo Santos Menezes, Carlos Leonardo Souza Cardoso, Fabio Andre Machado Porto

arXiv 2608.30114首次发表:更新:

发表机构

Laboratório Nacional de Computação Científica (LNCC)(国家科学计算实验室(LNCC))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对巴西葡萄牙语开源语言模型复现难、缺乏不确定性度量对比的问题,发布17.2亿参数的Manacá-1B模型及可复现流程,揭示分词器转换的评估陷阱并提供修复方案,所有资源可复现论文结果。

AI 中文摘要

巴西葡萄牙语在开源语言模型领域仍未得到充分覆盖,现有的少数模型难以复现,且常缺乏不确定性度量的对比。我们发布Manacá-1B,这是一个从头开始训练的、拥有17.2亿参数的仅解码器开源模型,专为巴西葡萄牙语设计,配备完全容器化、可复现的训练流程。预训练过程稳定,无跳过步骤或NaN(非数值)步骤,且能自恢复损失峰值,我们还发布了其完整日志与动态数据。我们在统一框架下,于四个葡萄牙语基准测试中将该模型与九个开源基线模型进行对比。所有对比均报告标准误差与配对显著性检验结果,且该框架已通过已发表数据验证。在末词预测任务中,Manacá-1B是7B规模以下最强的模型,在LAMBADA-PT数据集上以显著配对优势超越Tucano-1b1和Tucano-2b4;在常识补全任务中表现具有竞争力,在多项选择推理任务中接近随机水平,这与所有小型基础模型的表现一致。研究过程中,我们记录了一个具体的评估陷阱:将带有大小写折叠归一化的SentencePiece分词器转换为HuggingFace fast格式时,会静默丢失归一化器,导致所有大写标记被路由到字节回退机制,进而压低分数,而这在聚合指标中不可见。未修正的分词器使LAMBADA-PT准确率从45.3降至25.0;我们量化了该影响并提供了一行修复代码,可完全复现训练时使用的分词器。论文发布了代码、原始训练与评估日志、逐示例预测向量、模型权重及修正后的分词器,确保论文中的所有数值均可复现。

英文摘要

Brazilian Portuguese remains under-served by open language models, and the few that exist are difficult to reproduce and are often compared without measures of uncertainty. We release Manacá-1B, an open decoder-only model of 1.72 billion parameters trained from scratch for Brazilian Portuguese with a fully containerized, reproducible pipeline. The pretraining is stable, with zero skipped or NaN steps and self-recovering loss spikes, and we release its full log and dynamics. We evaluate the model against nine open baselines on four Portuguese benchmarks under a single harness. Every comparison reports a standard error and a paired significance test, and the harness is validated against previously published numbers. On last-word prediction Manacá-1B is the strongest model below the 7B scale, exceeding both Tucano-1b1 and Tucano-2b4 on LAMBADA-PT with large paired margins; it is competitive on commonsense completion and near chance on multiple-choice reasoning, as are all small base models. Along the way we document a concrete evaluation pitfall: converting a SentencePiece tokenizer with case-folding normalization to the HuggingFace fast format silently drops the normalizer, routing every capitalized token to byte-fallback and depressing scores in a way that is invisible in aggregate metrics. The uncorrected tokenizer lowered LAMBADA-PT accuracy from 45.3 to 25.0; we quantify the effect and provide a one-line fix that reproduces the training tokenizer exactly. Code, raw training and evaluation logs, per-example prediction vectors, the model weights, and the corrected tokenizer are released so that every number in this paper can be recomputed.

CommentsPreprint. Code: https://github.com/Instituto-IA-LNCC/manaca-1b-base ; model weights: https://huggingface.co/menezesbruno/manaca-1b-base

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑