arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ShriNep@EEUCA 2026:RAKSHAK - 用于毒性意图分类的具有原理蒸馏和拼图增强训练的多任务DeBERTa

ShriNep@EEUCA 2026: RAKSHAK - Multi-Task DeBERTa with Rationale Distillation and Jigsaw-Augmented Training for Toxic Intent Classification

Binayak Karki, Aryan Kafle, Pingala Ghimire

arXiv 2607.20450首次发表:更新:

发表机构

Mechi Multiple Campus; Northern Kentucky University; Himalaya College of Engineering(梅奇多校区; 北肯塔基大学; 喜马拉雅工程学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对GameTox共享任务中《坦克世界》聊天话语毒性意图分类难题,提出多任务DeBERTa - v3 - base框架RAKSHAK,结合原理蒸馏等方法,通过跨域转移和LLM生成样本增强训练数据,相比另一系统表现更优,多任务架构和跨域转移贡献显著。

AI 中文摘要

本文介绍了针对2026年ACL的EEUCA研讨会上的GameTox共享任务的两个系统,该任务要求将《坦克世界》聊天话语分类为六个细粒度的毒性意图类别(标签0 - 5)。严重的类别不平衡、特定领域的多语言俚语以及稀有类别(如威胁(标签4,60个样本)和极端主义(标签5,24个样本))的数据极度稀缺,使得这成为一个具有挑战性的分类问题。我们的主要提交作品RAKSHAK是一个多任务DeBERTa - v3 - base框架,结合了来自Qwen2.5 - 14B的原理蒸馏、监督对比损失和专用的稀有类别二元头。RAKSHAK的训练数据通过从拼图毒性评论数据集(16,225个样本映射到标签1 - 4)的跨域转移和100个用于标签5的LLM生成的极端主义样本进行增强。我们的次要系统(M1)在原始GameTox数据加上相同的100个极端主义样本上使用焦点损失对DeBERTa - v3 - base进行微调,不进行拼图转移。RAKSHAK在官方测试集上的宏F1为0.5883,在35个参赛团队中排名第7,而M1的宏F1为0.5252。对有和没有拼图数据的M1进行的对比消融表明,跨域转移贡献了 +2.6个F1点,而RAKSHAK的多任务架构又贡献了 +3.7个点。

英文摘要

This paper presents two systems for the GameTox Shared Task at the Workshop on EEUCA at ACL 2026, which requires classifying World of Tanks chat utterances into six fine-grained toxic intent categories (Labels 0-5). Severe class imbalance, domain-specific multilingual slang, and extremely scarce data for rare categories such as Threats (Label 4, 60 samples) and Extremism (Label 5, 24 samples) make this a challenging classification problem. Our primary submission, RAKSHAK (rak s. aka, Sanskrit for "Protector"), is a multi-task DeBERTa-v3-base (He et al., 2022) framework combining rationale distillation from Qwen2.5-14B (An et al., 2024), Supervised Contrastive Loss, and dedicated rare-class binary heads. RAKSHAK's training data is augmented with cross-domain transfer from the Jigsaw Toxic Comment dataset (16,225 samples mapped to Labels 1-4) and 100 LLM-generated extremism samples for Label 5. Our secondary system (M1) fine-tunes DeBERTa-v3-base with Focal Loss on the original GameTox data plus the same 100 extremism samples, without Jigsaw transfer. RAKSHAK achieves a Macro F1 of 0.5883 on the official test set, ranking 7th out of 35 participating teams, while M1 achieves 0.5252 Macro F1. An ablation comparing M1 with and without Jigsaw data shows that cross-domain transfer accounts for +2.6 F1 points, while RAKSHAK's multi-task architecture contributes a further +3.7 points.

Comments8 pages, 1 figure, EEUCA, ACL 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑