arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

解构刻板印象:面向有效多语言反言论的范围条件生成

Deconstructing Stereotypes: Scope-Conditioned Generation for Effective Multilingual Counterspeech

Greta Damo, Elias Urios Alacreu, Elena Cabrio, Paolo Rosso, Serena Villata

arXiv 2609.16906首次发表:更新:

发表机构

Université Côte d’Azur; CNRS; INRIA; I3S; PRHLT Research Center, Universitat Politècnica de València; ValgrAI Valencian Graduate School and Research Network of Artificial Intelligence(蔚蓝海岸大学; 法国国家科学研究中心; 法国国家信息与自动化研究所; I3S实验室; 瓦伦西亚理工大学PRHLT研究中心; ValgrAI瓦伦西亚人工智能研究生院与研究网络)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有反言论生成方法回复泛化、未针对隐性刻板印象的问题,提出范围条件生成框架,将结构化刻板印象特征融入大模型提示,在英意西三语人工数据集上显著提升事实性、具体性、说服力与有效性。

AI 中文摘要

反言论(CS)——利用推理和替代观点直接回应网络仇恨言论(HS)的回应——已成为内容移除的一种替代方案。然而,当前的自动反言论生成方法经常产生泛泛而谈、无效的回复,未能针对仇恨言论背后的隐性刻板印象。为弥合这一差距,我们提出了一种新颖的范围条件生成框架,该框架将结构化的刻板印象特征明确整合到大型语言模型提示中。我们在一个新颖的、人工策划的、以英语、意大利语和西班牙语标注的数据集上验证了我们的方法。大量评估表明,刻板印象条件提示在所有三种语言中均显著优于通用基线,在事实性、具体性、说服力和有效性方面,对显性和隐性刻板印象均取得了显著提升。

英文摘要

Counterspeech (CS) - direct responses that counter online Hate Speech (HS) using reasoning and alternative viewpoints - has emerged as an alternative to content removal. Current automatic CS generation methods, however, frequently produce generic, ineffective replies that fail to target the implicit stereotypes behind HS. To bridge this gap, we propose a novel scope-conditioned generation framework that explicitly integrates structured stereotype characteristics into Large Language Models prompts. We validate our approach on a novel, human-curated dataset annotated in English, Italian, and Spanish. Extensive evaluations show that stereotype-conditioned prompting substantially outperforms generic baselines across all three languages, obtaining significant gains in factuality, specificity, cogency, and effectiveness for both explicit and implicit implied stereotypes.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑