大锤还是手术刀?用于隐性仇恨言论检测的细粒度自适应框架
Sledgehammer or Scalpel? A Fine-grained Adaptive Framework for Implicit Hate Speech
- School of Information and Control Engineering, China University of Mining and Technology(中国矿业大学信息与控制工程学院)
- Yangzhou University(扬州大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对隐性仇恨言论检测中现有方法推理流程单一、计算冗余的问题,提出FAID框架,先细粒度分类再自适应处理不同类型样本,在四个基准数据集上显著优于SOTA基线。
AI中文摘要:
与带有明显亵渎性的显性攻击不同,隐性仇恨言论通过隐喻和语境暗示将恶意隐藏在看似顺从的表达中,这使得在线内容审核中的检测工作颇具挑战性。尽管现有的基于预训练语言模型(PLM)或大语言模型(LLM)的方法表现良好,但它们通常对所有样本采用单一推理流程,忽略了细粒度的语言差异,还会为较简单的案例造成不必要的计算开销。我们发现在线仇恨言论并非单一整体,而是呈现多种形式,因此定义了三个细粒度类别:浅层型、目标型和语境依赖型。据此,我们提出了细粒度自适应隐性仇恨言论检测框架(FAID),这是一种新颖的框架,它先进行细粒度分类,再针对特定类别进行自适应调整。具体而言,对于具有表层可识别意图的浅层样本,该框架采用轻量提示微调以实现快速分类;对于将恶意意图与隐藏目标绑定的目标型评论,我们设计了知识增强方法,以迭代优化模型并揭示隐藏目标;对于缺乏背景信息的语境依赖型评论,我们利用智能体框架自动生成提示以拓展语境、推断缺失的背景信息并识别模糊的恶意意图。这种自适应架构将计算资源集中于复杂的隐性样本,同时避免对浅层样本进行冗余推理。在四个基准数据集上的实验表明,FAID的性能显著优于现有最优(SOTA)基线方法。
英文摘要:
Unlike explicit attacks with obvious profanity, implicit hate speech hides malice within seemingly compliant expressions through metaphors and contextual hints, making its detection in online content review challenging. While existing PLM- or LLM-based methods perform well, they typically apply a single reasoning process to all samples. This overlooks fine-grained linguistic nuances and causes unnecessary computation for simpler cases. We observe that online hate speech is not monolithic but manifests in varied forms. We therefore define three fine-grained categories: Shallow, Targeted, and Context-Dependent. Accordingly, we propose Fine-grained Adaptive Implicit Hate speech Detection (FAID), a novel framework that first performs fine-grained classification and then adapts to specific categories. Specifically, for Shallow samples with surface-identifiable intents, the framework adopts lightweight prompt-tuning for rapid classification; for Targeted comments that bind malicious intent to concealed targets, we design knowledge augmentation to iteratively refine the model and reveal hidden targets; for Context-Dependent comments lacking background information, we utilize an agentic framework that automatically generates prompts to evolve context, infer missing background information and identify ambiguous malicious intents. This adaptive architecture focuses computational resources on complex implicit samples while avoiding redundant reasoning for shallow samples. Experiments on four benchmark datasets demonstrate that FAID significantly outperforms SOTA baselines.