arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用整数线性规划提升混合话语的语言识别性能

Improving Language Identification for Code-Switched Utterances with Integer Linear Programming

Joanna Radoła, Josep Maria Crego, François Yvon

arXiv 2609.05099首次发表:更新:

发表机构

Sorbonne Université; SYSTRAN by ChapsVision(索邦大学; ChapsVision旗下SYSTRAN)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对混合话语的语言识别挑战,改进SOTA方法MaskLID,将其优化算法重构为整数线性规划,在10种语言实验中显著提升混合基准性能并发布代码数据。

AI 中文摘要

自动识别混合(CS)话语仍是语言识别(LID)系统面临的挑战,这类文本在大语言模型的训练数据中代表性不足。本文重新探讨无训练、可检测任意语言组合的混合识别SOTA方法MaskLID,完成三项核心贡献:(a)揭示并解决MaskLID过度依赖词级语言关联分数的关键问题;(b)将底层优化算法重新表述为整数线性规划,使我们可尝试大量清晰且可解释的约束条件;(c)上述改进大幅提升基线系统性能,在涉及10种不同语言的实验中,我们观察到混合基准测试的性能显著提升。我们发布代码与数据以保障可复现性。

英文摘要

Automatic identification of code-switched (CS) utterances remains a challenge for language identification (LID) systems, causing such texts to be underrepresented in the training data of Large Language Models. In this paper, we revisit MaskLID, a state-of-the art approach for CS identification, which requires no training and detects arbitrary language combinations. We make three main contributions: (a) we reveal, and address, a major issue of MaskLID: its overreliance on word-level language association scores; (b) we reformulate the underlying optimization algorithm as an Integer Linear Program, enabling us to experiment with a large set of clear and interpretable constraints; (c) each of these improvements vastly improves the baseline system, as we illustrate in experiments involving 10~diverse languages, where we observe a strong boost in performance on CS benchmarks. We release our code and data for reproducibility.

CommentsAccepted to Findings of EMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑