发表机构
ETSI de Telecomunicación, Universidad Politécnica de Madrid(马德里理工大学电信工程技术学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对EEUCA 2026游戏聊天毒性分类任务,提出结合紧凑型变压器集成与语言信息中介的三阶段管道,通过特定策略解决类不平衡问题,在官方测试集上取得优异成绩,管道具领域可移植性。
AI 中文摘要
本文描述了我们针对EEUCA 2026游戏聊天毒性分类共享任务的系统。我们实现了一个三阶段管道,将两个紧凑型变压器(DeBERTa-v3-base,184M;XLM-RoBERTa-base,278M)的集成与一个语言信息中介(LIM)相结合,该中介通过语料库支持的词汇规范化、类条件一元语法评分、多语言亵渎检测以及基于言语行为理论的代理目标分析来解决模型间的分歧。LIM专门针对少数类(仇恨与骚扰、威胁和极端主义),这是现实世界游戏审核中最关键的安全类别。为解决极端类不平衡(无毒与极端主义比例为1450:1),我们仅使用提供的训练数据引入了两阶段数据增强策略。我们的系统在官方测试集上实现了0.6441的宏F1和0.9062的准确率,在所有团队中宏F1排名第三,准确率排名第一。所提出的管道具有领域可移植性:适应其他游戏平台只需替换特定游戏的实体词典。代码可在这个https URL\_EEUCA上公开获取。
英文摘要
This paper describes our system for the EEUCA 2026 Shared Task on toxicity classification in gaming chat. We implement a three-stage pipeline combining an ensemble of two compact transformers (DeBERTa-v3-base, 184M; XLM-RoBERTa-base, 278M) with a Linguistically-Informed Mediator (LIM) that resolves inter-model disagreements through corpus-backed lexical normalization, class-conditional unigram scoring, multilingual profanity detection, and agentive targeting analysis grounded in speech act theory. The LIM specifically targets the minority classes (Hate \& Harassment, Threats, and Extremism), which are the most safety-critical categories in real-world gaming moderation. To address the extreme class imbalance (1{,}450:1 Non-toxic to Extremism ratio), we introduce a two-stage data augmentation strategy using only the provided training data. Our system achieves a Macro F1 of 0.6441 and accuracy of 0.9062 on the official test set, ranking 3rd in Macro F1 and 1st in accuracy among all teams. The proposed pipeline is domain-portable: adapting to other gaming platforms requires substituting only the game-specific entity lexicon. Code is publicly available at https://github.com/Anmol2059/thaulab\_EEUCA.