利用上下文词表示进行形态丰富语言的句法分析的多任务学习
Multi-Task Learning by using Contextualized Word Representations for Syntactic Parsing of a Morphologically Rich Language
浏览论文内容
中文总结 AI 辅助
本文针对乌尔都语这一形态丰富语言,提出结合上下文词表示的多任务学习框架,通过统一序列标注和共享架构同时学习成分与依存结构,显著提升分析性能,F1达91.39,依存得分85.69。
中文摘要 AI 辅助
我们解决了乌尔都语(一种形态丰富的语言)的句法分析挑战,并展示了成分分析和依存分析的最新成果。本文做出了四项主要贡献:1)通过开发语言特定的中心词和短语到依存标签的映射规则,将CLE-UTB短语结构树库转换为依存树库;2)一种新颖的序列标注方案,将分析任务转化为统一表示;3)在从网络收集的220百万词元的乌尔都语大型语料库上训练上下文词表示;4)使用两种学习范式(单任务学习和多任务学习)开发分析框架。应用了多种后处理规则以提高自动转换的依存结构树库的质量。所提出的序列标注方案使得共享架构能够同时从两种语法结构中学习句法结构,从而改善泛化能力。实验表明,多任务学习设置显著提升了分析性能,成分分析达到了91.39的F1分数(提高了3.29分),依存分析达到了85.69的标注附着分数(提高了1.49分)。这些结果表明,学习跨任务表示带来了可衡量的收益,并推动了乌尔都语句法分析的最新进展。
英文摘要
We address the challenge of syntactic parsing for Urdu, a morphologically rich language, and present state-of-the-art results for both constituency and dependency parsing. This paper offers four major contributions: 1) the conversion of the CLE-UTB phrase structure treebank into a dependency treebank by developing language-specific head-word and phrase-to-dependency label mapping rules; 2) a novel sequence labeling scheme that transforms the parsing task into a unified representation; 3) the training of contextualized word representations on a large 220 million tokens Urdu corpus collected from the web; and 4) development of parsing framework using two learning paradigms, single-task and multi-task learning. Several post-processing rules are applied to improve the quality of the automatically converted dependency structure treebank. The proposed sequence labeling scheme enables the use of a shared architecture that learns the syntactic structures from both grammatical structures simultaneously and hence improves generalization. Experiments show that the multi-task learning setup significantly enhances parsing performance, achieving an F1 score of 91.39 for constituency parsing (an improvement of 3.29 points) and a labeled attachment score of 85.69 for dependency parsing (an improvement of 1.49 points). These results demonstrate that learning cross-task representations provides measurable benefits and advances the state of syntactic parsing for Urdu.
发表机构
- VTT Technical Research Centre of Finland Ltd.(芬兰VTT技术研究中心有限公司)
- University of Konstanz(康斯坦茨大学)
- Al-Khawarizmi Institute of Computer Science(花剌子密计算机科学研究所)
- University of Engineering and Technology(工程技术大学)
- Umm Al-Qura University(乌姆·库拉大学)
- Copenhagen University(哥本哈根大学)
机构由 AI 辅助整理,请以论文原文为准。