修正即标注:为中世纪拉丁语文献引导训练依存句法分析器
Correction as Annotation: Bootstrapping a Dependency Parser for Documentary Medieval Latin
浏览论文内容
中文总结 AI 辅助
针对中世纪拉丁语文献缺乏可用解析工具的问题,提出以专家修正模型预标注作为训练数据的方法,显著提升解析性能并减少数据需求。
中文摘要 AI 辅助
中世纪文献资料在现有自然语言处理工具中仍未得到充分支持。在1258年至1446年间马赛编纂的160份财产清单集合上,五个现成的拉丁语树库模型均未达到可用性能。最佳标注附着得分为0.62,最佳形态感知得分为0.24。性能与体裁或时期邻近性均不相关。为解决这一不足,我们利用这些不完善模型作为副产品生成了领域内训练数据。在九次迭代中,每次迭代中一个模型预先标注200个句子;专家修正标注;修正后的句子用于训练后续模型,批次独立于模型状态采样,未采用主动学习选择。对1,804个句子投入33小时的标注工作,将通用词性标注准确率从0.80提升至0.98,标注附着从0.48提升至0.92,在报告指标上优于所有基线,同时使用的训练数据比最大基线少97%。标注者工作量从占词元的54%下降至14-18%的稳定水平,这是一个无需单独金标准且可作为停止标准的操作进度指标。
英文摘要
Medieval documentary sources remain inadequately served by existing natural language processing tools. None of the five readily available Latin treebank models attains usable performance on a collection of 160 inventories compiled in Marseille between 1258 and 1446. The best labelled attachment score is 0.62 and the best morphology-aware score is 0.24. Performance does not correlate with either genre or period proximity. To address this shortfall, in-domain training data was generated as a by-product of using these inadequate models. In each of nine iterations, a model pre-annotated 200 sentences; an expert corrected the annotations; and the corrected sentences were used to train the subsequent model, with batches sampled independently of model state, without active-learning selection. Thirty-three hours of annotation effort over 1,804 sentences increased universal part-of-speech accuracy from 0.80 to 0.98 and labelled attachment from 0.48 to 0.92, outperforming all baselines on the reported metrics while using 97% less training data than the largest one of them. Annotator effort declined from 54% of tokens to a plateau of 14-18%, an operational progress metric that requires no separate gold standard and can serve as a stopping criterion.
发表机构
- Harvard University(哈佛大学)
机构由 AI 辅助整理,请以论文原文为准。