arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.21988cs.CL

分析语言模型中的自我伤害表征:跨架构研究

Analysing Self-Harm Representations in Language Models: a Cross-Architecture Study

  • Cardiff University(卡迪夫大学)

机构由 AI 辅助整理,请以论文原文为准。

Luis Espinosa-Anke, Carla Perez-Almendros

AI总结:

研究使用自然语言处理技术检测自我伤害内容的挑战,通过对两个数据集和四个模型进行两项实验,分析语言模型对自我伤害内容的表征,发现自我伤害信息在网络层特定深度结晶,且最准确探针不一定最线性可分,Gemma - 3 - 4B表征方式独特。

AI中文摘要:

使用自然语言处理技术检测自我伤害内容具有挑战性,且是高风险任务,需要高精度以实现及时干预或标记有风险用户。本文分析语言模型如何表征此类内容,其在自我伤害检测、语言模型干预、治理和监管中有下游应用。聚焦两个数据集和四个模型,进行两项主要实验:一是在两个自我伤害数据集上训练和评估各模型所有层的线性探针,发现自我伤害信息在网络层最后3 - 7%(93至97%深度)结晶;二是提取对比性自我伤害方向,发现最准确的探针不一定是最线性可分的,Gemma - 3 - 4B以略有不同、更复杂的方式表征这种对比性自我伤害方向。

英文摘要:

Self-harm content is particularly challenging to detect using NLP techniques, and is also a high-stakes task which requires the highest accuracy to enable timely intervention or flagging at-risk users. We therefore present an analysis of how LLMs represent such self-harm content, which has downstream applications in self-harm detection, LLM intervention and governance and policing. In this paper, we focus on two datasets and four models, and perform two main experiments: (1) We train and evaluate linear probes across all layers of each model on two self-harm datasets: X-Sensitive and SH-Detection. Across both corpora, self-harm information crystallizes in the final 3 - 7% of network layers (93 to 97% depth). (2) We extract contrastive self-harm directions and, after performing a normaliation step, we find that the most accurate probes are not necessarily the most linearly separable. In particular, we find Gemma-3-4B to represent this \textit{contrastive self-harm direction} in a slightly different, more intricate way than the other LLMs.

↑