用于分层攻击性语言检测的级联与联合建模
Cascading versus Joint Modeling for Hierarchical Offensive Language Detection
浏览论文内容
中文总结 AI 辅助
研究细粒度攻击性语言检测中级联与联合建模范式,提出三级级联检测系统及验证机制,通过实验对比两种范式在准确性、参数数量和推理延迟方面的表现,明确了级联架构准确性优势与其部署成本间的权衡。
中文摘要 AI 辅助
细粒度攻击性语言检测将标签组织成层次结构,存在级联分解和联合多任务建模两种建模范式。以往工作很少对这两种范式在准确性、参数数量和推理延迟方面进行直接、可控的比较,也很少验证所选的类不平衡处理策略是否最优。本文提出了一个三级级联检测系统,其训练策略是针对每个子任务定制的,还有两种验证机制。实验表明,级联系统在官方测试集的三个子任务上分别获得了0.795、0.716和0.557的宏F1分数。消融研究表明,仅凭不平衡严重程度直觉配置损失函数是次优的;根据消融结果重新配置可提高性能和稳定性。相对于联合多任务模型,级联架构在所有三个子任务上都实现了更高的准确性,在最严重不平衡的子任务上宏F1增益为7.1分,代价是参数增加三倍和推理延迟增加1.67倍。这些结果在级联架构的准确性优势与其部署成本之间建立了明确、可量化的权衡。
英文摘要
Fine-grained offensive language detection organizes labels into a hierarchical structure, for which two modeling paradigms exist: cascaded decomposition and joint multi-task modeling. Prior work rarely provides a direct, controlled comparison of the two paradigms in terms of accuracy, parameter count, and inference latency, and rarely verifies whether a chosen class-imbalance handling strategy is actually optimal. This paper proposes a three-level cascaded detection system whose training strategy is customized per subtask, together with two verification mechanisms. First, a controlled ablation study determines the best class-imbalance handling strategy for each subtask. Second, a joint multi-task model with a shared encoder is trained as an architectural control, yielding real measurements along the dimensions of accuracy, parameter count, and inference latency. Experiments show that the cascaded system attains macro-F1 scores of 0.795, 0.716, and 0.557 on the three subtasks of the official test set. The ablation study reveals that configuring the loss function purely by imbalance-severity intuition is suboptimal; reconfiguring based on the ablation results improves both performance and stability. End-to-end cascade evaluation shows that roughly one-fifth of the errors in the cascade pipeline originate from the first-stage filter and cannot be corrected by subsequent stages. Relative to the joint multi-task model, the cascaded architecture achieves higher accuracy on all three subtasks, with a 7.1-point macro-F1 gain on the most severely imbalanced subtask, at the cost of three times the parameters and 1.67 times the inference latency. Together, these results establish an explicit, quantifiable trade-off between the accuracy advantage of cascaded architectures and their deployment cost.
发表机构
- School of Electronic and Information Engineering, Beijing Jiaotong University(北京交通大学电子信息工程学院)
- National Computer Network Emergency Response Technical Team/Coordination Center of China (CNCERT/CC)(国家计算机网络应急技术处理协调中心)
机构由 AI 辅助整理,请以论文原文为准。