发表机构
Uppsala University; Lund University(乌普萨拉大学; 隆德大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出4.05亿参数的Stoicheia字符级掩码扩散模型,通过多维度输入设计实现古希腊语文本恢复、句法分析等任务,在三项实验中均优于对照模型,且在Ithaca测试集上显著降低字符错误率、提升Top-1准确率。
AI 中文摘要
我们推出Stoicheia,这是一个拥有4.05亿参数的古希腊语字符级掩码扩散编码器,其输入可分解为五个对齐且独立可掩码的维度:字母、词与句子边界、变音符号、大小写以及标点。因此,单一主干模型无需针对特定任务重新分词,即可完成缺损文本的恢复、重新分词、添加重音与标点。我们在一个开放的、经修订锁定的3.8亿词语料库上对其进行预训练,并发布11个检查点:10个旋转且去污染的折,保证对于任何给定的文学段落,至少有一个发布模型从未见过该文本;另有1个未接触过文献文本的检查点。三项实验——受损铭文与纸莎草纸的重建、形态句法标注与依存句法分析,以及带韵律扫描的长音符号标注——各配备了随机初始化的对照模型,以分离字符级扩散预训练的贡献:铭文重建的字符错误率(CER)提升5.6个点,句法分析的标签附件分数(LAS)提升12.9个点,长音符号标注的平衡准确率提升6.0个点。在Ithaca自身的测试拆分上,使用相同的冻结样本与严格评分,Stoicheia将字符错误率相对于现有最优系统从24.6(Ithaca)和23.5(其2025年Aeneas框架后继系统)降至15.5,并将Top-1准确率从63.0和64.0提升至74.5。
英文摘要
We introduce Stoicheia, a 405M-parameter character-level masked-diffusion encoder for Ancient Greek whose input factors into five aligned, independently maskable planes: letters, word and sentence boundaries, diacritics, capitalization, and punctuation. A single backbone can therefore restore lacunae, re-segment, accentuate, and punctuate unspaced text without task-specific retokenization. We pretrain it on an open, revision-pinned corpus of 380M words and release eleven checkpoints: ten rotated, decontaminated folds, guaranteeing that for any given literary passage at least one released model has never seen its text, and one with no exposure to documentary texts. Three experiments - reconstruction of damaged inscriptions and papyri, morphosyntactic tagging and dependency parsing, and macronization with metrical scansion - each carry a matched random-initialization control, isolating what character-level diffusion pretraining contributes: 5.6 CER points on inscription reconstruction, 12.9 LAS on parsing, and 6.0 points of balanced accuracy on macronization. On Ithaca's own test split, with identical frozen samples and strict scoring, Stoicheia reduces character error relative to both prior state-of-the-art systems, from 24.6 (Ithaca) and 23.5 (its 2025 Aeneas-framework successor) to 15.5, and raises top-1 accuracy from 63.0 and 64.0 to 74.5.
Comments12 pages, 7 tables. Models, datasets and code released: https://huggingface.co/collections/Ericu950/stoicheia-6a6fbf9800c82d93020a7ceb and https://github.com/ericu9500/stoicheia