arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

更大的上下文窗口,更少的过度修正:面向最小编辑语法纠错的提示与批处理优化

Larger Context Window, Fewer Overcorrections: Optimizing Prompts and Batching for Minimal-Edit Grammatical Error Correction

Kateryna Karpo, Artem Chernodub

arXiv 2609.10810首次发表:更新:

发表机构

Ukrainian Catholic University; YouScan; Zendesk(乌克兰天主教大学; YouScan; Zendesk)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对最小编辑语法纠错中LLM过度修正问题,提出基于分类法指令、句子批处理和LLM辅助提示优化的方法,在BEA-2019上达到F0.5=78.32,接近微调SOTA。

AI 中文摘要

最小编辑语法纠错(GEC)对于零样本和少样本提示的大型语言模型(LLMs)而言是一项具有挑战性的任务,这些模型会系统性地过度修正,并通过重写格式良好的片段来降低$F_{0.5}$分数。虽然微调提供了有效的解决方案,但它带来了大量的基础设施需求。我们提出了一种基于提示的方法,通过在GEC提示方法学上的三项进展,缩小了与微调模型之间的差距。首先,我们引入了基于分类法的指令,通过一份全面的语法错误规则列表来强制执行最小编辑约束,使LLM具备一个有界的、与指标对齐的可修正编辑范围,这对最强的模型有益,但总体上仍依赖于模型。其次,我们证明将多个未修正的句子批处理到单个输入上下文中,可作为针对过度修正的定向正则化器,系统地降低不同LLM家族的编辑率;我们假设这源于自注意力分数的有限容量引起的注意力稀释效应。最后,LLM辅助的提示优化进一步完善了这些指令。由Gemini 3.1-Pro驱动,我们的提示在BEA-2019测试集上达到了$F_{0.5}=78.32$,确立了新的基于提示的最先进水平(SOTA),同时将差距缩小到与微调单模型SOTA(Staruch等人,2025)仅0.38分。代码、提示和输出均已公开。

英文摘要

Minimal-edit Grammatical Error Correction (GEC) is a challenging task for zero- and few-shot prompted Large Language Models (LLMs), which systematically overcorrect and degrade $F_{0.5}$ by rewriting well-formed spans. While fine-tuning provides an effective solution, it imposes substantial infrastructure demands. We introduce a prompt-based approach that closes the gap to fine-tuned models through three advances in GEC prompting methodology. First, we introduce taxonomy-based instructions to enforce minimal-edit constraints with a comprehensive list of grammatical error rules, equipping the LLM with a bounded, metric-aligned scope of correctable edits, which benefits the strongest models while remaining model-dependent overall. Second, we show that batching multiple uncorrected sentences into a single input context acts as a targeted regularizer against overcorrection, systematically reducing the edit rate across diverse LLM families; we hypothesize this arises from attention dilution effect induced by the bounded capacity of self-attention scores. Finally, LLM-assisted Prompt Optimization refines these instructions. Powered by Gemini 3.1-Pro, our prompt achieves $F_{0.5}=78.32$ on the BEA-2019 test set - establishing a new prompt-based SOTA while shrinking the gap to the fine-tuned single-model SOTA (Staruch et al., 2025) to a mere $0.38$ points. Code, prompts, and outputs are publicly available.

CommentsAccepted for publication at EMNLP 2026 (Findings)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑