通过信念自蒸馏进行用户模型提取
User Model Extraction via Belief Self-Distillation
浏览论文内容
中文总结 AI 辅助
本文提出信念自蒸馏(BSD)框架,从冻结LLM中提取可读且可因果写入的用户表征,实现比隐藏状态引导更强的干预,并发现跨模型共享的用户表征几何结构,对AI安全有重要启示。
中文摘要 AI 辅助
大型语言模型(LLMs)会隐式推断其用户的属性,并相应地调整自身行为,然而这些信念仍然难以检查并进行因果操控。我们引入了信念自蒸馏(Belief Self-Distillation, BSD),这是一个统一的读写框架,通过学习一个紧凑的用户表征来桥接线性探针与因果探针,该表征既可以解码,也可以被写回模型中。冻结的LLM充当自身的教师,在没有外部标注的情况下从自然对话中蒸馏信念。与传统探针不同,BSD不仅分离出激活中存在的信息,还分离出可直接测试其因果作用的内部状态。在多个模型家族中,BSD能够忠实地恢复用户信念,并且能够实现比匹配的隐藏状态引导(hidden-state steering)显著更强的干预效果。至关重要的是,我们发现拒绝(refusal)不仅取决于请求本身,还取决于模型推断出的用户意图:在保持请求不变的情况下,改变这一信念会改变拒绝行为。我们进一步发现了一个显著的跨模型规律:独立训练的LLM在表征其用户时收敛于一个共享的几何结构。综合来看,这些结果揭示了隐式用户模型是可读且可因果写入的内部状态,对AI安全具有直接影响,塑造了模型如何根据其认为正在交互的对象来调整安全决策。
英文摘要
Large language models (LLMs) implicitly infer attributes of their users and adapt their behavior accordingly, yet these beliefs remain difficult to inspect and causally manipulate. We introduce Belief Self-Distillation (BSD), a unified read-write framework that bridges linear and causal probing by learning a compact user representation that can be both decoded and written back into the model. The frozen LLM acts as its own teacher, distilling beliefs from natural conversations without external annotations. Unlike conventional probing, BSD isolates not only information present in activations, but a state whose causal role can be directly tested. Across multiple model families, BSD faithfully recovers user beliefs and enables substantially stronger interventions than matched hidden-state steering. Crucially, we find that refusal depends not only on the request, but on the model's inferred user intent: changing this belief alters refusal while holding the request fixed. We further uncover a striking cross-model regularity: independently trained LLMs converge on a shared geometry for representing their users. Together, these results reveal implicit user models as readable and causally writable internal states with direct implications for AI safety, shaping how models condition safety decisions on whom they believe they are interacting with.
发表机构
- Technical University of Munich(慕尼黑工业大学)
- Helmholtz Zentrum München(亥姆霍兹慕尼黑中心)
机构由 AI 辅助整理,请以论文原文为准。