arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

借词还是语码转换?标注边界而非模型驱动的哈萨克语-俄语语码转换识别

Loanword or Switch? The Annotation Boundary, Not the Model, Drives Kazakh-Russian Code-Switching Identification

Bogdan Savelyev

arXiv 2608.00581首次发表:更新:

AI 中文总结

该研究针对哈萨克语-俄语语码转换识别中借词易被误判的问题,构建了规范标注的黄金LID数据集,发现性能瓶颈源于标注边界而非模型,为相关识别任务提供了关键依据。

AI 中文摘要

现成的LID工具和字母启发式规则会将哈萨克语-俄语社交文本过度标注为混合文本:哈萨克语中的俄语借词在共用西里尔字母体系下看起来像语码转换。我们发布了一个文档级黄金LID数据集,其指南将整合式借词保留为哈萨克语,仅将子句级转换标注为混合,还提供了LID后用于过滤优先级联的仅混合情感池。在共享LID测试中,FastText、Lingua、原始及窗口化HeLI、字符三元组NB、XLM-R的性能从弱到强不等,差距表明瓶颈在于借词与转换的标注边界,而非仅模型类别。

英文摘要

Off-the-shelf LID and letter heuristics over-label Kazakh-Russian social text as mixed: Russian loanwords inside Kazakh look like code-switching under a shared Cyrillic script. We release a document-level gold LID set whose guideline keeps integrated borrowings as Kazakh and reserves mixed for clause-level switches, plus a mixed-only sentiment pool used after LID in a filter-first cascade. On a shared LID test, FastText, Lingua, raw and windowed HeLI, character-trigram NB, and XLM-R range from weak to strong performance. The gap shows the bottleneck is the loanword-vs-switch annotation boundary, not model class alone.

Comments5 pages. Preprint. Submitted to W-NUT 2026. Code and data: https://github.com/naadgob/KazNLP

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑