arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.14178cs.CLcs.CYcs.SI

一种用于有效反叙事生成与优化的多阶段智能体框架

A Multi-Stage Agentic Framework for Effective Counter-Narrative Generation and Refinement

发表机构特拉维夫大学 · 早稻田大学
查看机构详情
  • Tel Aviv University(特拉维夫大学)
  • Waseda University(早稻田大学)

机构由 AI 辅助整理,请以论文原文为准。

Carmel Kronfeld, Sharva Gogawale, Tetsuro Kobayashi, Irad Ben-Gal

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出多阶段智能体框架,通过试点实验和多智能体迭代优化生成反叙事,经人类验证和自动化安全分析,证明其能降低亲俄叙事强度并优于基线,为反仇恨言论提供可扩展方案。

中文摘要 AI 辅助

社交媒体上仇恨言论和错误信息的快速传播对民主社会构成了挑战,因为直接压制可能加深两极分化、加剧公众不信任并强化极端主义叙事。基于大语言模型的反叙事提供了一种有前景的方式来降低这些风险,但其有效性取决于修辞和风格选择,而这些选择目前仍缺乏深入理解。我们提出了一种多阶段基于智能体的框架,用于生成、优化和评估反叙事,并将其应用于关于乌克兰战争的亲俄仇恨和错误信息叙事,同时该框架可适应其他领域。一项由人类评估者参与的试点实验识别出有效的技巧-风格组合,例如重复与情感框架相结合能增强说服力。基于这些见解,我们引入了一个多智能体优化过程,该过程迭代地改进反叙事,以提升其说服力、情感参与度和可分享性。在人类验证确认改进后,一项自动化安全分析表明,我们优化后的反叙事在质量上达到或超过了专家撰写的反言论。随后的一项模拟实验显示,这些反叙事降低了亲俄叙事的感知强度,并始终优于普通大语言模型基线,这凸显了一条针对仇恨言论和错误信息进行可扩展、叙事特定干预的路径。本工作的代码和数据可在该 https URL 公开获取。

英文摘要

The rapid diffusion of hate speech and misinformation on social networks challenges democratic societies, since direct suppression efforts may deepen polarization, fuel public distrusts, and strengthen extremist narratives. LLM-driven counter-narratives (CNs) offer a promising way to reduce those risks, yet their effectiveness depends on rhetorical and stylistic choices that remain poorly understood. We present a multi-stage agent-based framework for generating, refining, and evaluating CNs, applied to pro-Russian hate and misinformation narratives on the war with Ukraine and adaptable to other domains. A pilot experiment with human evaluators identifies effective technique style pairings, such as repetition with emotional framing enhancing persuasiveness. Building on these insights, we introduce a multi-agent refinement process that iteratively improves CNs for persuasiveness, emotional engagement, and shareability. After human validation confirmed improvement, an automated safety analysis shows that our refined CNs match or improve on expert-written counterspeech. A simulated experiment then shows that they reduce the perceived strength of pro-Russian narratives and consistently outperform a vanilla LLM baseline, highlighting a pathway toward scalable, narrative-specific interventions against hate speech and misinformation. Code and data accompanying this work are publicly available at https://github.com/carmelkron/inlg2026-counter-narratives.

↑