发表机构
Kyushu University(九州大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对濒危语言伊良部语语符标注成本高的问题,用小型BiLSTM-CRF模型构建神经标注管道。发现黄金词性标注可提升语法语符标注,词性层级能减少所需语符化数据量,虽全自动管道未完全实现增益,但为文献实践提供四线标注建议。
AI 中文摘要
话语数据是田野语言学中语法编写的主要实证基础,但生成语符化文本成本高昂,每分钟录音大约需要一小时的工作量。对于濒危语言,与母语者核实分析的时间有限,因此自动化部分语符化工作流程具有直接的文献价值。我们使用精心设计的小型透明双向长短期记忆条件随机场(BiLSTM-CRF)模型,为伊良部琉球语实现了一个完整的神经标注管道(词素分割、词性标注、语符标注),并在一个现实的严格约束下进行评估:大约一小时的完全标注话语作为全部监督资源。对标注本身的两个因素进行了操作:其丰富程度(有无词性层级)和数量(训练预算从6到47分钟)。黄金词性标注将语法语符标注提高了4.4(标准差0.7)分(在所有5个种子中均显著),并且随着数据量减少,增益增加(数据量为四分之一时增加11.6分);词性层级使达到给定准确率所需的语符化数据量减少一半以上。在全自动管道中,这种增益尚未实现:标注器仍有12%的词素出错,错误的词性标注比没有词性标注更能误导语符标注模型。价值是潜在的而非丧失的:用可控噪声降低黄金词性标注显示,随着标注器准确率提高,增益会恢复,在接近我们标注器当前88%的准确率时达到收支平衡,在92%-96%的准确率时恢复1.6到3.2分。我们最后为文献实践提出了一个具体建议:进行四线标注——文本、词性、语符、翻译。
英文摘要
Discourse data are the primary empirical basis of grammar writing in field linguistics, but producing interlinearized text is notoriously expensive - on the order of one hour of work per minute of recording. For endangered languages, where the time remaining to verify analyses with native speakers is itself limited, automating parts of the interlinearization workflow has direct documentary value. We implement a full neural annotation pipeline (morpheme segmentation, POS tagging, glossing) for Irabu Ryukyuan using deliberately small, transparent BiLSTM-CRF models, and evaluate it under a realistic hard constraint: approximately one hour of fully annotated discourse as the entire supervised resource. Two factors of the annotation itself are manipulated: its richness (with or without a POS tier) and its quantity (training budgets from 6 to 47 minutes). Gold POS improves grammatical glossing by +4.4 (SD 0.7) points (significant in all 5 seeds), and the gain grows as data shrink (+11.6 points at a quarter of the data); a POS tier more than halves the amount of glossed data needed to reach a given accuracy. In a fully automatic pipeline this gain is not yet realized: the tagger still errs on 12% of morphemes, and an incorrect POS misleads the glossing model more than no POS at all. The value is latent rather than lost: degrading gold POS with controlled noise shows the gain returning as tagger accuracy rises, with break-even near our tagger's current 88% and +1.6 to +3.2 points recovered at 92-96%. We conclude with a concrete recommendation for documentation practice: annotate quadrilinearly - text, POS, gloss, translation.
Comments15 pages, 9 figures, 8 tables. Code, corpus, and complete per-sentence test outputs: https://github.com/MichinoriShimoji/ML_autogloss