发表机构
University of Arizona(亚利桑那大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出句法感知位置嵌入SiPE,将句法先验注入Transformer,使模型在SyntaxGym、GLUE基准分别提升10.3%、8.2%,困惑度降9.0%,在句法监督与推理成本间建立新帕累托前沿。
AI 中文摘要
Transformer中的位置嵌入(PE)用于编码 token 的距离和顺序,但在很大程度上与句法结构无关。我们提出句法感知位置嵌入(Syntax-informed Positional Embeddings, SiPE),该方法在预训练期间从依存句法分析中学习轻量级句法先验,并将其注入到三大主流位置嵌入族(绝对、相对、旋转)中,适用于编码器和解码器,同时不改变自注意力和模型的其余架构。我们确定了该先验应进入模型的位置和方式,发现这取决于架构:对于使用相对位置嵌入的自回归解码器,当该先验与注意力分数的相对位置项相乘结合时效果最强,优于将其注入输入嵌入、注入自注意力或同时注入位置和注意力项的情况;而对于编码器,最佳方式是直接将其添加到输入嵌入中,与编码器的原生位置机制结合。我们发现,使用 SiPE 预训练的模型在 SyntaxGym 基准测试中性能提升高达 10.3%,同时与没有句法监督的基础模型相比,困惑度降低了 9.0%——这一指标几乎所有现有的句法注入方法都会使其下降。关键的是,这些增益不仅限于句法泛化:SiPE 还能改善现实世界的语言理解,在 GLUE 基准测试中,其得分比未使用 SiPE 训练的模型提高了高达 8.2%。与现有在推理时对多个句法分析结果取平均或在运行时丢弃句法的句法语言模型不同,SiPE 基于单一的句法分析结果进行条件设置,在句法监督和推理成本之间建立了新的帕累托前沿。
英文摘要
Positional embeddings (PE) in Transformers encode token distance and order but are largely agnostic to \textit{syntactic structure}. We introduce \textbf{S}yntax-\textbf{i}nformed \textbf{P}ositional \textbf{E}mbeddings (\textbf{SiPE}), which learns a lightweight syntactic prior from dependency parses during pretraining and injects it across all three dominant PE families (absolute, relative, rotary), for both encoders and decoders, leaving self-attention and the rest of the architecture untouched. We isolate \emph{where} and \emph{how} the prior should enter the model, and find it depends on the architecture: for autoregressive decoders that use relative PE, the prior is strongest when coupled multiplicatively with the relative-position term of the attention score, outperforming injection into the input embeddings, into self-attention, or into the positional and attention terms jointly---while for encoders it is best added directly to the input embeddings, composing with each encoder's native positional mechanism. We find that models pre-trained with SiPE improve on the SyntaxGym benchmark by up to $10.3\%$ while simultaneously reducing perplexity by $9.0\%$ over a base model with no syntactic supervision---a metric nearly every existing syntax-injection method instead degrades. Crucially, these gains extend beyond syntactic generalization: SiPE also improves real-world language understanding, raising scores on the GLUE benchmark by up to $8.2\%$ over a model trained without it. Unlike existing syntactic language models that marginalize over many parses at inference or discard syntax at runtime, SiPE conditions on a single parse, establishing a new Pareto frontier between syntactic supervision and inference cost.
Comments21 pages, 9 figures