arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

古希腊元音长度的自动标注

Automatic Annotation of Ancient Greek Vowel Length

Albin Thörn Cleland, Eric Cullhed

arXiv 2608.01935首次发表:更新:

发表机构

Lund University; Uppsala University(隆德大学; 乌普萨拉大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对古希腊语缺乏大规模公开长音标注语料库的问题,构建了首个通用长音标注器,其输出可训练小型字符级Transformer提升韵律NLP任务效果。

AI 中文摘要

古希腊自然语言处理(NLP)领域的现有研究依赖的语料库未对α、ι、υ(统称双元音dichrona)的音位元音长度进行歧义消解。根据词素、形态、连音、句法,以及时期、体裁、诗体形式的惯例,这些字母中的每一个都可表示长元音或短元音。确定并标注正确长度的过程被称为“长音标注(macronizing)”,由于单词形式数量庞大且每个实例存在语境依赖性,这是一个长尾问题。目前尚无大规模公开可用的古希腊长音标注语料库,因此需要一个独立的长音标注器。尽管此前研究已展示如何构建静态的、适配特定语料库的元音长度词典,本文则构建了首个可处理任意古希腊语文本的通用长音标注器。该标注器接收带有词元、词性和形态标注的标准CoNLL-U格式输入,通过一组递归模块,使不太常见的单词形式能继承同一词素更常见形式的标注。该长音标注器的主要应用是生成机器学习的训练数据:研究表明,基于该标注器输出训练的小型字符级Transformer,能在规则系统未标注的案例上实现泛化,在诗体和散文的人工标注黄金基准上的准确率与规则系统相当或更高。研究还显示,长音标注可改善下游韵律NLP任务,如诗行扫描。

英文摘要

Prior work in Ancient Greek NLP relies on corpora that do not disambiguate the phonemic vowel length of alpha, iota, and ypsilon, together known as the dichrona. Depending on lexeme, morphology, sandhi, syntax, and conventions of period, genre, and verse form, each of these letters can represent either a long or a short vowel. Deciding and marking the correct length is known as "macronizing", a long-tail problem given the sheer mass of word forms and the context dependency of individual instances. No macronized corpus of Ancient Greek is publicly available at scale, so a stand-alone macronizer is needed. While previous work has shown how to build a static, corpus-bespoke vowel-length dictionary, the present paper constructs the first general-purpose macronizer for arbitrary Ancient Greek input. Given input carrying lemma, part-of-speech, and morphological annotation in the standard CoNLL-U format, a set of recursive modules lets less common word forms inherit markup from more common forms of the same lexical word. The macronizer's chief application is generating training data for machine learning: we show that a small character-level transformer trained on the macronizer's own output learns to generalize past the cases the rule-based system leaves unmarked, matching or exceeding its accuracy on a gold-standard, manually annotated benchmark of verse and prose. We also show that macronization can improve downstream prosodical NLP tasks like verse scansion.

Comments5 pages, 0 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑