HatePrototypes: Interpretable and Transferable Representations for Implicit and Explicit Hate Speech Detection
HatePrototypes: 用于隐性与显性仇恨言论检测的可解释且可迁移的表示
机构 * Laboratoire Hubert Curien, UMR CNRS 5516(于贝尔·居里实验室,法国国家科学研究中心5516联合研究单位) ; Université Lumière Lyon 2(里昂第二大学) ; Université Claude Bernard Lyon 1(克洛德·贝尔纳里昂第一大学) ; ERIC, Lyon(里昂ERIC实验室) ; École Centrale de Lyon(里昂中央理工学院) ; LIRIS CNRS UMR 5205(LIRIS实验室,法国国家科学研究中心5205联合研究单位)
专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI
AI总结 本文提出HatePrototypes,通过语言模型优化的类级向量表示,实现显性和隐性仇恨言论的跨任务迁移,无需重复微调。
Journal ref In Proceedings of the Fifteenth Language Resources and Evaluation Conference, pages 4387-4399, 2026