发表机构
University of Chinese Academy of Sciences(中国科学院大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对有害聊天对话不断变化的问题,提出结合有序推理链(ORC)正则化的BRACE模型,在多领域多危害类型检测中取得优异性能,各组件均有贡献。
AI 中文摘要
有害聊天对话会通过类型转换和词汇规避不断变化,但研究发现它们存在不变的原则,即有序推理链(ORC),包含反复出现的话题、有害语言指标、严重程度层次和类型特征,可帮助捕捉频繁变化的词汇表达中的关键信息。本文提出BRACE,将ORC编码为四个可微阶段(话题→指标→严重程度→类型)并引入中间监督,作为结合直接头的结构化正则化器,同时结合基于原型的特征增强和特征路径解缠。评估结果显示,在4个领域和5类危害中,BRACE的危害类型宏F1值为0.934(基于RoBERTa-wwm-ext,3个随机种子的均值),使用解码器主干(Qwen3-1.7B LoRA)时达到0.949。 ablation研究表明所有组件均对BRACE有贡献,ORC的结构分解使其能区分具有语义歧义的危害类型。免责声明:本文可能包含令部分读者不适的内容。
英文摘要
Harmful chat dialogues are ever-shifting through type-shifting and lexical evasion, yet we find they share invariant principles, i.e., an Ordered Reasoning Chain (ORC) of recurring topics, harm language indicators, severity hierarchies, and type characteristics, which can help us capture the key information in the frequently changing lexical expressions. We propose BRACE, which encodes the ORC as four differentiable stages (Topic -> Indicator -> Severity -> Type) with intermediate supervision, serving as a structured regularizer blended with direct heads, and supported by prototype-based feature augmentation and feature path disentanglement. The evaluation results show that, across 4 domains and 5 harm categories, BRACE achieves harm-type macro F1 of 0.934 (RoBERTa-wwm-ext, 3-seed mean), with decoder backbones (Qwen3-1.7B LoRA) reaching 0.949. Ablation studies show that all components contribute to BRACE, and the structural decomposition of ORC enables BRACE to distinguish harmful types with semantic ambiguity. Disclaimer: This paper may contain content that is disturbing to some readers.
Comments9 pages, 4 figures, conference