发表机构
School of Foreign Languages, Peking University(北京大学外国语学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出STAGEET分阶段类型化编辑标记框架,以阿拉伯语为案例,在QALB-2014等数据集上实现了语法纠错的最优性能,同时提升了修正过程的可解释性。
AI 中文摘要
序列到编辑(Seq2Edit)方法通过预测输入的编辑标签而非生成完整修正句,使语法纠错(GEC)高效且具备局部可解释性。然而其可解释性主要是操作性的:标签仅指定字符串的变更方式,单一编辑词汇表无法始终明确修正类型。本文提出STAGEET,一种分阶段类型化编辑标记框架,将Seq2Edit的监督信号重组为类型化可执行阶段,并将编辑操作扩展至修正类别。STAGEET将修正分解为有序的中等粒度类型化阶段序列;每个阶段从自身标签空间预测,改写当前假设一次,并将得到的中间句子传递给下一阶段。我们将该框架实例化为两种模型:一种是带阶段特定适配器的端到端共享编码器多头模型,另一种是每个阶段配备独立标记器的完全专用变体。在QALB-2014和ZAEBUC数据集上的实验表明,感知类别的分阶段修正在保持具有竞争力的基于编辑的GEC性能的同时,呈现出更具可检查性的修正轨迹,并在QALB-2014数据集上取得了当前最优(SOTA)结果。
英文摘要
Sequence-to-edit approaches make grammatical error correction (GEC) efficient and locally interpretable by predicting edit labels over the input rather than generating a full corrected sentence. Their interpretability, however, is primarily operational: a label specifies how the string should change, but a single edit vocabulary does not always reveal the type of correction being made. We propose STAGEET, a stage-wise typed edit-tagging framework that reorganizes Seq2Edit supervision into typed executable stages and extends edit operations to correction categories. STAGEET decomposes correction into an ordered sequence of medium-grained typed stages; each stage predicts from its own label space, rewrites the current hypothesis once, and passes the resulting intermediate sentence to the next stage. We instantiate the framework as both an end-to-end shared-encoder multi-head model with stage-specific adapters and a fully specialized variant with one independent tagger per stage. Experiments on QALB-2014 and ZAEBUC show that category-aware staged correction retains competitive edit-based GEC performance while exposing a more inspectable correction trajectory, and attains state-of-the-art results on QALB-2014.