arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向中文学习者语法错误标注的分层分类体系

A Layered Taxonomy for Chinese Learner Grammatical Error Annotation

Mengyang Qiu, Jungyeul Park

arXiv 2609.02153首次发表:更新:

发表机构

Saint Elizabeth University; KAIST(圣伊丽莎白大学; 韩国科学技术院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出一种分层分类体系,链接CGEC与教学错误分析,通过覆盖度分析和大语言模型一致性研究验证其有效性,用于中文学习者语法错误标注。

AI 中文摘要

中文学习者写作中的语法错误标注需要兼具一致性和语言学意义的标签。本文提出了一种将计算中文语法错误纠正(CGEC)与教学错误分析相链接的分层方案。该方案首先识别字符和标点层面的正字法错误,通过编辑操作和子类型进行标注;其他错误则采用三层核心标签,结合编辑操作、语言学领域和词性,还可附加针对中文特有的体貌、情态、比较、论元结构和补语的扩展。该分类体系基于CGEC资源、学习者错误分类和普通话语法,通过对自动提取的MuCGEC编辑的覆盖度分析,以及五项大语言模型将其应用于样本的初步一致性研究进行评估。结果支持该分层方法,同时指出需要进一步完善的类别边界。

英文摘要

Grammatical error annotation in Chinese learner writing requires labels that are both consistent and linguistically meaningful. This paper proposes a layered scheme linking computational Chinese grammatical error correction (CGEC) with pedagogical error analysis. The scheme first identifies character- and punctuation-level orthographic errors, labeling them by edit operation and subtype. Other errors receive a three-layer core label combining edit operation, linguistic domain, and part of speech, with optional Chinese-specific extensions for aspect, modality, comparison, argument structure, and complements. Drawing on CGEC resources, learner-error taxonomies, and Mandarin grammar, the taxonomy is evaluated through a coverage analysis of automatically extracted MuCGEC edits and a preliminary consistency study in which five large language models apply it to a sample. The results support the layered approach while identifying category boundaries requiring further refinement.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑