arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.18156cs.CL

TeochewBench:一个经人工审校的潮汕话汉字翻译基准

TeochewBench: A Human-Reviewed Benchmark for Teochew Hanzi Translation

Jianan Wu

首次发表
浏览论文内容

中文总结 AI 辅助

提出经人工审校的潮汕话汉字翻译基准TeochewBench,含300条表达,评估13个模型共7800条预测,发现高特异性表达得分低且跨模型差异小,Qwen3.5-27B最优。

中文摘要 AI 辅助

潮汕话拥有庞大的使用群体,并展现出独特的词汇、句法和语用特征,然而用于评估大语言模型的文本资源仍然有限。我们提出了TeochewBench,一个经人工审校的基准,包含300条潮汕话汉字表达,用于评估从潮汕话汉字到普通话和英语的翻译。该数据集涵盖五个类别:基础词汇;日常句子;潮汕话特有表达;语气、礼貌与语境;以及习语、歧义和文化特定表达。一位主要的潮汕话母语审校者逐条检查了所有条目并按需进行了修订,另有两位潮汕话使用者核验了部分条目。我们的主要评估涵盖11个官方通用后训练模型在审校后的数据集上的两个翻译方向,共产生6,600条预测。两个官方基础检查点提供了1,200条预测用于补充诊断,使模型总数达到13个,预测总数达到7,800条。我们还加入了一个汉字复制对照组,该组直接返回源输入不变,以评估共享汉字如何影响翻译成普通话的自动评分。Qwen3.5-27B在评估的检查点中取得了最高的总体chrF风格分数,为60.63,其次是Qwen2.5-72B-Instruct(56.61)、Gemma-3-27B-IT(56.36)和GLM-4-32B-0414(55.82)。在11个主要评估模型中,平均chrF风格分数从低特异性条目的69.25下降到高特异性条目的27.52。高特异性表达得分较低,且跨模型差异较小,表明它们构成了所评估模型家族中共同的低分区域。汉字复制对照组进一步表明,低特异性条目中的表面重叠会显著影响翻译成普通话的自动评分。

英文摘要

Teochew has a substantial speaker community and exhibits distinctive lexical, syntactic, and pragmatic features, yet textual resources for evaluating large language models remain limited. We present TeochewBench, a human-reviewed benchmark comprising 300 Teochew Hanzi expressions for evaluating translation from Teochew Hanzi into Mandarin Chinese and English. The dataset covers five categories: basic vocabulary; everyday sentences; Teochew-specific expressions; tone, politeness, and context; and idiomatic, ambiguous, and culturally specific expressions. A primary Teochew-speaking reviewer examined all entries individually and revised them as needed, while two additional Teochew speakers verified selected items. Our main evaluation covers 11 official general-purpose post-trained models on the reviewed dataset in both translation directions, yielding 6,600 predictions. Two official base checkpoints provide 1,200 predictions for supplementary diagnostics, bringing the total to 13 models and 7,800 predictions. We additionally include a Hanzi-copy control, which returns the source input unchanged, to assess how shared Hanzi affect automatic scores for translation into Mandarin Chinese. Qwen3.5-27B achieved the highest overall chrF-style score among the evaluated checkpoints, at 60.63, followed by Qwen2.5-72B-Instruct at 56.61, Gemma-3-27B-IT at 56.36, and GLM-4-32B-0414 at 55.82. Across the 11 main-evaluation models, the mean chrF-style score decreased from 69.25 for low-specificity items to 27.52 for high-specificity items. High-specificity expressions received lower scores and exhibited smaller cross-model differences, suggesting that they constitute a shared low-scoring region across the model families evaluated here. The Hanzi-copy control further indicates that surface overlap in low-specificity items can substantially affect automatic scores for translation into Mandarin Chinese.

发表机构

  • Peking University(北京大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑