arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.10939cs.CLcs.AI

使用小型语言模型的低成本多语言短文本分类路由流水线

A Cost-Efficient Routing Pipeline for Multilingual Short-Text Classification Using Small Language Models

Wajdi Ben Saad, Safa Madiouni

AI总结:

本研究提出一种使用小型语言模型的多语言短文本分类路由流水线,通过选择性翻译低资源语言提升分类性能,在两个基准上验证了该策略的有效性。

AI中文摘要:

多语言短文本分类支持内容审核、客户支持路由、意图识别等业务系统,但综合评估往往掩盖了高资源语言与低资源语言之间的巨大差异。统一推理策略部署简单,但假设所有语言都能得到同等良好的支持。本研究评估一种固定列表路由策略:将较强语言保留在直接多语言路径中,选择性地将较弱语言翻译为英文后再进行零样本分类。该流水线完全自托管,使用预训练的紧凑句子编码器,无需任务特定的微调。我们在两个规模和标签粒度不同的基准上测试该方法:SIB-200的15语言子集,用于七类主题分类;MASSIVE的15区域子集,用于官方60意图清单上的意图分类。在SIB-200上,最佳整体配置为R1,仅翻译低资源层级:高、中层级的Macro-F1保持不变,而低层级Macro-F1从0.4632升至0.6828。在MASSIVE子集上,相同的低资源干预使低层级Macro-F1从0.2143升至0.4417,但最佳整体结果通过全翻译(R3)获得,Macro-F1为0.4647。在这两个基准中,选择性翻译是对较弱语言的可靠干预,而最优路由边界取决于任务。因此,我们报告通过层级级质量增益和层级级延迟进行路由,而非单一全局效率分数。

英文摘要:

Multilingual short-text classification supports operational systems such as content moderation, customer support routing, and intent recognition, yet aggregate evaluation often hides large differences between high-resource and low-resource languages. Uniform inference policies are simple to deploy, but they assume that all languages are equally well served. In this work, we evaluate a fixed-list routing strategy that keeps stronger languages on a direct multilingual path and selectively sends weaker languages through translation into English before zero-shot classification. The pipeline is fully self-hosted, uses pretrained compact sentence encoders, and requires no task-specific fine-tuning. We test the approach on two benchmarks chosen to differ in scale and label granularity: a 15-language subset of SIB-200 for seven-way topic classification and a 15-locale subset of MASSIVE for intent classification over an official 60-intent inventory. On SIB-200, the best overall configuration is R1, which translates only the low-resource tier: high-tier and mid-tier Macro-F1 remain unchanged, while low-tier Macro-F1 rises from 0.4632 to 0.6828. On the MASSIVE subset, the same low-tier intervention raises low-tier Macro-F1 from 0.2143 to 0.4417, but the best overall result is obtained by full translation, R3, at Macro-F1 0.4647. Across these two benchmarks, selective translation is a reliable intervention for weaker languages, whereas the optimal routing boundary depends on the task. We therefore report routing through tier-level quality gains and tier-level latency rather than a single global efficiency score.

补充信息

↑