结合证据的回应!用于仇恨类别感知反仇恨言论生成的多智能体内存高效推理框架
Counter with Evidence! A Multi-Agent Memory Efficient Reasoning Framework for Hate Category Informed Counterspeech Generation
浏览论文内容
中文总结 AI 辅助
该研究针对现有反仇恨言论生成未区分仇恨言论类别的问题,提出多智能体框架FIRE并构建数据集FactualCS,实验显示FIRE效果优于基线且毒性更低。
中文摘要 AI 辅助
反仇恨言论可有效削弱网络仇恨言论的影响。尽管现有研究探索了自动化反仇恨言论生成,但大多侧重风格控制,将仇恨言论视为同质内容,忽视了不同形式的滥用行为需要截然不同的反仇恨言论策略。为解决这一缺口,我们提出FIRE(Factuality Informed Multi-Agent Reasoning Framework,事实性感知多智能体推理框架),该框架首先将仇恨言论划分为五大类之一(错误信息、刻板印象、阴谋论、非人化、非事实性),再将其映射至针对性的反仇恨言论风格。为支撑FIRE,我们构建了新数据集FactualCS,包含4784个实例,提供仇恨类别、推理轨迹及证据映射的标注,这些是现有研究缺失的、对 grounded 生成至关重要的元素。对28种基线配置的综合评估显示,FIRE尽管使用规模小于20亿参数的紧凑智能体,仍显著优于现有方法;相较于最强基线,FIRE在事实性准确率和类别特定准确率上分别提升约12%和11%,同时毒性降低约11%。进一步人工评估确认,FIRE生成的回应显著优于最强基线,凸显其实际部署的有效性。这些发现表明,剖析仇恨言论的潜在意图对生成安全、有效且语境精准的反仇恨言论至关重要。
英文摘要
Counterspeech effectively neutralizes the impact of online hate. Although prior work explores automated counterspeech generation, it largely emphasizes stylistic control while treating hate speech as homogeneous, overlooking that distinct forms of abuse require fundamentally different counterspeech strategies. To address this gap, we introduce FIRE (Factuality Informed Multi-Agent Reasoning Framework) that first decomposes hate speech into one of the five distinct categories (misinformation, stereotype, conspiracy, dehumanizing, non-factual), and then maps it to a targeted counterspeech style. To facilitate FIRE, we curate FactualCS, a novel dataset of $4,784$ instances that provides the annotations regarding hate categories, reasoning traces, and evidence mappings, which are critical elements for grounded generation that are missing in prior work. A comprehensive evaluation across $28$ baseline configurations demonstrates that FIRE significantly surpasses existing methods, despite using compact agents ($<$2B). FIRE achieves a $\sim$ $12 \%$ and $\sim$ $11 \%$ improvements in factual and category-specific accuracy respectively, while simultaneously reducing toxicity by $\sim$ $11 \%$ relative to the strongest baselines. Further human evaluation confirms that responses generated by FIRE are significantly preferred over the strongest baselines, underscoring its effectiveness for real-world deployment. These findings show that decomposing the underlying intent of hate speech is essential for generating safe, effective, and contextually precise counterspeech.
发表机构
- Indian Institute of Technology Delhi(印度德里理工学院)
机构由 AI 辅助整理,请以论文原文为准。