POSPAN:用于语言模型预训练的位置约束跨度掩码
POSPAN: Position-Constrained Span Masking for Language Model Pre-training
- JD AI Research(京东人工智能研究院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
POSPAN提出位置约束跨度掩码框架,统一现有方法,通过结合跨度长度与位置约束分布,在NLU基准上优于仅长度方法和原始MLM,并给出理论解释。
AI中文摘要:
跨度级掩码语言建模(MLM)已被证明比原始的单令牌MLM对预训练语言模型更有利,因为实体/短语及其依赖关系对语言理解至关重要。以往的工作仅考虑具有某些离散分布的跨度长度,而忽略了跨度之间的依赖关系,即假设掩码跨度的位置是均匀分布的。在本文中,我们提出了POSPAN,一个通用框架,通过跨度长度分布和位置约束分布的组合,允许多样化的位置约束跨度掩码策略,该框架统一了所有现有的跨度级掩码方法。为了验证POSPAN在预训练中的有效性,我们在多个NLU基准数据集上对其进行了评估。实验结果表明,位置约束能够广泛增强跨度级掩码,并且我们最佳的POSPAN设置始终优于仅考虑跨度长度的对应方法和原始MLM。我们还对掩码语言模型中的位置约束进行了理论分析,以阐明POSPAN为何有效的原因,证明了POSPAN的合理性和必要性。
英文摘要:
Span-level masked language modeling (MLM) has shown to be advantageous to pre-trained language models over the original single-token MLM, as entities/phrases and their dependencies are critical to language understanding. Previous works only consider span length with some discrete distributions, while the dependencies among spans are ignored, i.e., assuming that the positions of masked spans are uniformly distributed. In this paper, we present POSPAN, a general framework to allow diverse position-constrained span masking strategies via the combination of span length distribution and position constraint distribution, which unifies all existing span-level masking methods. To verify the effectiveness of POSPAN in pre-training, we evaluate it on the datasets from several NLU benchmarks. Experimental results indicate that the position constraint is capable of enhancing span-level masking broadly, and our best POSPAN setting consistently outperforms its span-length-only counterparts and vanilla MLM. We also conduct theoretical analysis for the position constraint in masked language models to shed light on the reason why POSPAN works well, demonstrating the rationality and necessity of POSPAN.