发表机构
School of Computation, Information and Technology, TU Munich; Munich Center for Machine Learning; Munich Data Science Institute(慕尼黑工业大学计算、信息与技术学院; 慕尼黑机器学习中心; 慕尼黑数据科学研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对LLMs在仇恨言论检测等敏感领域表现不佳的问题,基于Qwen3,通过统一36个英语仇恨言论数据集进行指令调优,提升了模型在有害内容缓解上的跨域、跨语言泛化性能。
AI 中文摘要
大语言模型(LLMs)在广泛的通用自然语言处理任务中展现出令人印象深刻的性能;然而,它们在仇恨言论检测等敏感领域的有效性仍不明确。此前将提示式LLMs与基于编码器的最先进模型(例如BERT变体(Roy等人,2023;Dönmez等人,2024))进行比较的研究仅显示出微小的提升,表明LLMs可能在仇恨言论检测或缓解方面并不出色。在本研究中,我们从指令调优的角度重新审视该问题。通过全面统一涵盖多种标注方案的36个英语仇恨言论数据集,我们基于Qwen3(Qwen团队,2025)对一个通用LLM进行微调,专门用于仇恨言论缓解。我们的结果不仅在域内基准上展现出最先进的性能,还在跨域和跨语言泛化方面取得了显著提升——而这些正是基于编码器的专用分类器常遇到困难的领域。
英文摘要
Large language models (LLMs) demonstrate impressive performance across a wide range of general NLP tasks; however, their effectiveness in sensitive domains, such as hate speech detection, remains less clear. Prior studies comparing prompted LLMs with state-of-the-art encoder-based models (e.g., BERT variants (Roy et al., 2023; Dönmez et al., 2024)) have shown only marginal gains, suggesting that LLMs may not excel in hate speech detection or mitigation. In this work, we revisit this question through the lens of instruction tuning. By thoroughly unifying 36 English hate speech datasets spanning multiple labeling schemes, we fine-tune a generalist LLM, based on Qwen3 (Qwen Team, 2025), specifically for hate speech mitigation. Our results demonstrate not only state-of-the-art performance on in-domain benchmarks but also substantial improvements in cross-domain and cross-lingual generalization--areas where encoder-based specialist classifiers often struggle.