arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于渐进式微调序列到序列Transformer的泰米尔语上下文拼写与语法修正

Contextual Tamil Spelling and Grammar Correction Using Progressively Fine-Tuned Sequence-to-Sequence Transformers

Karthikeyan A, Jaya Nirmala S, Sangeetha Sivanesan, Indhu R, Pranav Kumar, Bharat Jude Johnson, Vishnu Ram

arXiv 2609.03273首次发表:更新:

发表机构

National Institute of Technology, Tiruchirappalli; Tamil University, Thanjavur(国家技术研究所蒂鲁吉拉伯利分校; 泰米尔大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对泰米尔语低资源特性及现有方法无法处理上下文错误的问题,提出端到端序列到序列模型,通过四阶段渐进式微调mT5-small和mBART-50,在合成语料库上取得连音准确率87.5%、主谓一致准确率43.5%的结果。

AI 中文摘要

泰米尔语的拼写和语法修正颇具挑战性,因为它是一种黏着性低资源语言,拥有丰富的动词形态、复杂的词边界连音(语音转换)规则,以及包含247个不同字母的文字系统。现有研究多采用基于规则的方法、统计n元语法模型、最小编辑距离,或结合Transformer重排序器的混合流水线,针对的是词级表面错误;这类方法无法可靠处理需要句子级理解的上下文错误,例如主谓一致、时态一致性或跨词连音。我们提出一种端到端的序列到序列架构,并在最多包含657720对带噪-干净泰米尔语句子的合成语料库上微调mT5-small和mBART-50,该语料库涵盖10类错误。两种模型主干均遵循相同的四阶段渐进式训练计划,每个阶段针对一类弱点:表面噪声(v2)、上下文语法(v3)、单处连音(v4)、多处跨词连音(v5)。在与所有训练数据不重叠的1000句平衡诊断集上,我们的最优模型mBART-50 v5达到69.3%的top-1精确匹配准确率,其中连音准确率为87.5%,主谓一致准确率为43.5%。该训练计划是性能提升的关键:引入上下文句子对后,主谓一致准确率从1.0%升至52.5%;引入多处连音句子对后,连音准确率从0%升至87.5%。我们还量化了该领域未报道过的精确率-召回率权衡:连音召回率的提升会以身份准确率的单调下降为代价。最后,Tamil-LLaMA-7B-Instruct的零样本准确率为19.0%,经3次演示后升至24.7%,而复制基线为20.0%,这表明适配泰米尔语的指令模型若缺乏特定任务监督,无法迁移至专业的句子级修正任务。

英文摘要

Tamil spell and grammar correction is challenging because Tamil is an agglutinative low-resource language with rich verbal morphology, complex sandhi (phonetic transformation) rules at word boundaries, and a script of 247 distinct letters. Prior work targets word-level surface errors with rule-based methods, statistical n-gram models, Minimum Edit Distance, or hybrid pipelines with a transformer re-ranker; such methods cannot reliably handle contextual errors - subject-verb agreement, tense consistency, or cross-word sandhi - which require sentence-level understanding. We propose an end-to-end sequence-to-sequence formulation and fine-tune mT5-small and mBART-50 on a synthetic corpus of up to 657,720 noisy-clean Tamil sentence pairs spanning ten error categories. Both backbones follow the same four-stage progressive schedule, each stage targeting one weakness: surface noise (v2), contextual grammar (v3), single-site sandhi (v4), and multi-site cross-word sandhi (v5). On a 1,000-sentence balanced diagnostic set verified disjoint from all training data, our best model, mBART-50 v5, reaches 69.3% top-1 exact-match accuracy, with 87.5% on sandhi and 43.5% on subject-verb agreement. The schedule is what produces these gains: subject-verb accuracy rises from 1.0% to 52.5% once contextual pairs are introduced, and sandhi from 0% to 87.5% once multi-site sandhi pairs are. We additionally quantify a precision-recall trade-off this literature has not reported: sandhi recall is paid for monotonically in identity accuracy. Finally, Tamil-LLaMA-7B-Instruct reaches 19.0% zero-shot and 24.7% with three demonstrations against a 20.0% copy baseline, showing that a Tamil-adapted instruction model does not transfer to specialised sentence-level correction without task-specific supervision.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑