发表机构
Sorbonne Université; SYSTRAN by ChapsVision(索邦大学; ChapsVision旗下SYSTRAN)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对混合话语的语言识别挑战,改进SOTA方法MaskLID,将其优化算法重构为整数线性规划,在10种语言实验中显著提升混合基准性能并发布代码数据。
AI 中文摘要
自动识别混合(CS)话语仍是语言识别(LID)系统面临的挑战,这类文本在大语言模型的训练数据中代表性不足。本文重新探讨无训练、可检测任意语言组合的混合识别SOTA方法MaskLID,完成三项核心贡献:(a)揭示并解决MaskLID过度依赖词级语言关联分数的关键问题;(b)将底层优化算法重新表述为整数线性规划,使我们可尝试大量清晰且可解释的约束条件;(c)上述改进大幅提升基线系统性能,在涉及10种不同语言的实验中,我们观察到混合基准测试的性能显著提升。我们发布代码与数据以保障可复现性。
英文摘要
Automatic identification of code-switched (CS) utterances remains a challenge for language identification (LID) systems, causing such texts to be underrepresented in the training data of Large Language Models. In this paper, we revisit MaskLID, a state-of-the art approach for CS identification, which requires no training and detects arbitrary language combinations. We make three main contributions: (a) we reveal, and address, a major issue of MaskLID: its overreliance on word-level language association scores; (b) we reformulate the underlying optimization algorithm as an Integer Linear Program, enabling us to experiment with a large set of clear and interpretable constraints; (c) each of these improvements vastly improves the baseline system, as we illustrate in experiments involving 10~diverse languages, where we observe a strong boost in performance on CS benchmarks. We release our code and data for reproducibility.
CommentsAccepted to Findings of EMNLP 2026