发表机构
Technische Hochschule Nürnberg Georg Simon Ohm(纽伦堡乔治·西蒙·欧姆应用技术大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对德语社交媒体有害内容检测的类别不平衡挑战,构建了含三个正交维度的九投票器集成系统,在2026年GermEval共享任务四个子任务中均获第一。
AI 中文摘要
德语社交媒体中的有害内容会造成现实世界危害,从煽动行动号召到刑事诽谤不等。2026年GermEval共享任务将其检测分为四个子任务,技术挑战在于严重的类别不平衡问题:有害类别占比稀少,且与占主导的多数类别共享表层语言,但其宏F1值是决定任务得分的关键。因此关键杠杆并非更强的单一模型,而是错误独立性。基于该洞见,构建了针对每个子任务的九投票器集成系统,涵盖大语言模型(LLM)、训练方法和类别范围三个正交维度,主要通过内部交叉验证进行选择。该系统在隐藏测试集上的宏F1值分别为:C2A(行动号召)89.56、DBO(诽谤)71.63、VIO(暴力)54.84、DEF(其他有害内容)83.02,在全部四个子任务中均排名第一。
英文摘要
Harmful content in German social media does real-world damage, from calls to action to criminal defamation. The GermEval 2026 shared task scores its detection in four subtasks. The technical challenge is a severe class imbalance. The harmful classes are rare and share surface language with the dominant majority class, yet under macro-F1 they decide the score. The decisive lever is then not a stronger single model but error independence. This insight becomes a per-subtask nine-voter ensemble spanning three orthogonal axes: LLM, training method and class scope. Selected mainly on internal cross-validation, the system reaches macro-F1 of 89.56 (C2A), 71.63 (DBO), 54.84 (VIO) and 83.02 (DEF) on the hidden test set, placing first on all four subtasks.
CommentsAccepted at the GermEval 2026 Shared Task on Harmful Content Detection @ KONVENS 2026 (1st place on all four subtasks)