arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.12361cs.CLcs.CY

新术语,新毒性:基于共识的新词毒性检测——通过搜索增强型大语言模型(LLMs)

New Terms, New Toxicity: Consensus-based Chinese Neologism Toxicity Detection via Search-Augmented LLMs

  • Tsinghua University(清华大学)
  • ICT, CAS(中国科学院计算技术研究所)
  • National Computer Network Emergency Response Technical Team Coordination Center of China(国家计算机网络应急技术处理协调中心)
  • IIE, CAS(中国科学院工程热物理研究所)
  • School of Cyber Security, UCAS(中国科学院大学网络空间安全学院)
  • JCSS, Tsinghua University(清华大学软件学院)
  • Science City (Guangzhou) Digital Technology Group Co., Ltd.(广州科学城数字科技集团有限公司)

机构由 AI 辅助整理,请以论文原文为准。

Shiyao Cui, QingLin Zhang, Di Wang, Yida Lu, Zhexin Zhang, Jinhua Gao, Jinglin Yang, Min He, Han Qiu, Minlie Huang

AI总结:

本文针对毒性新词的检测问题,提出捕捉毒性新词起源与共识标准的分类法,构建风险词表,引入搜索增强型框架SeTox,实验显示3B规模的SeTox性能优于近期大规模模型。

AI中文摘要:

新词是指在意义或形式上新兴的术语,可成为表达毒性的新载体,例如“country girl”是针对女性主义的污名化标签。这类毒性新词看似无害,却在公众共识中演变为毒性用法,对内容审核系统构成挑战,且相关研究仍不充分。本文研究如何检测通过新词表达的隐性毒性。我们首先提出一种分类法,用于捕捉毒性新词的起源和共识验证标准,随后构建涵盖广泛观察到的风险类别的词表。为捕捉基于公众共识的毒性,我们引入SeTox,这是一个搜索增强型框架,使静态大语言模型(LLMs)能够纳入实时网络上下文以进行新词毒性检测。实验表明,即使是30亿参数规模的模型,SeTox也优于近期的大规模模型,证明其可扩展性,可将现实世界知识纳入毒性新词检测。免责声明:本文包含冒犯性内容,可能会令部分读者不适。

英文摘要:

Neologisms, emerging terms in meaning or form, can serve as new vehicles for toxic expression, like "country girl" as a stigmatizing label targeting feminism. Such toxic neologisms appear benign but have evolved into toxic usage in public consensus, posing challenges to moderation systems and remaining underexplored. In this paper, we investigate how to detect implicit toxicity expressed via neologisms. We first propose a taxonomy that captures the origins and consensus-verification criteria of toxic neologisms, followed by the construction of a lexicon spanning widely observed risk categories. To capture toxicity grounded in public consensus, we introduce SeTox, a search-augmented framework that enables static large language models (LLMs) to incorporate real-time web context for neologism toxicity detection. Experiments show that SeTox, even with 3B-scale models, outperforms recent large-scale models, demonstrating its scalability to incorporate real-world knowledge for toxic neologism detection. Disclaimer: this paper has offensive contents that may be disturbing to some readers.

补充信息

↑