噪声出,偏见入:通过闭环激活引导在扩散语言模型中进行定向偏见注入
Noise Out, Bias In: Targeted Bias Injection in Diffusion Language Models via Closed-Loop Activation Steering
浏览论文内容
中文总结 AI 辅助
针对掩码扩散语言模型,提出基于闭环激活引导的定向偏见注入攻击,利用去噪轨迹控制通道,显著提升目标答案偏好,揭示需审计服务栈。
中文摘要 AI 辅助
掩码扩散语言模型(dLLMs)通过迭代去噪掩码位置来生成文本,在每个标记被确定之前多次重新预测该标记。自回归解码器在提交答案的那一步仅暴露一次答案的分布;而dLLM则在提交前的每个去噪步骤都暴露该分布,我们表明攻击者可以利用这一点。由于答案在多个去噪步骤中仍可被修改,能够访问内部激活的攻击者可以观察模型产生选定答案的可能性,并相应调整干预。基于这一观察,我们研究了定向偏见注入,这是一种将冻结的dLLM引导至攻击者选定的人口统计答案的攻击方法。该攻击使用一个简单的比例-积分(PI)控制器,在去噪过程中跟踪目标答案概率,并实时调整引导向量的强度。在模糊的BBQ问题(正确答案为弃权(不执行))上,我们的攻击将LLaDA-8B-Instruct对目标群体的偏好从1.8个百分点提高到16.7个百分点,是固定强度引导基线最强效果的三倍以上;在SocialStigmaQA上,它将污名化答案的选择从17.6%提高到58.1%。当适配到其他人口统计目标时,同一攻击可使答案偏移高达37个百分点,每次攻击在一张GPU上约需40分钟。对于主要目标,反馈是攻击起作用的关键:在标记提交步骤上以相同平均强度进行恒定引导所产生的偏移远小于此,同时破坏近三倍的输出,而针对每个示例单独设置的恒定强度仍远不及。我们的发现将去噪轨迹识别为dLLMs中一个新的控制通道,并呼吁进行偏见审计,检查服务栈而非仅检查冻结模型。
英文摘要
Masked diffusion language models (dLLMs) generate text by iteratively denoising masked positions, re-predicting each token multiple times before it is committed. An autoregressive decoder exposes an answer's distribution once, at the step that commits it; a dLLM exposes it at every denoising step before commitment, and we show that an adversary can exploit this. Since an answer remains open to revision over many denoising steps, an adversary with access to internal activations can watch how likely the model is to produce a chosen answer and adjust the intervention accordingly. Building on this observation, we study targeted bias injection, an attack that steers a frozen dLLM toward a demographic answer selected by the adversary. The attack uses a simple proportional-integral (PI) controller that tracks the target-answer probability during denoising and adapts the strength of a steering vector on the fly. On ambiguous BBQ questions where the correct answer is abstention, our attack raises LLaDA-8B-Instruct's preference for the targeted group from 1.8 to 16.7 percentage points, more than three times the strongest fixed-strength steering baseline, and on SocialStigmaQA it raises the selection of stigmatizing answers from 17.6% to 58.1%. Fitted to other demographic targets, the same attack shifts answers by up to 37 percentage points, and each attack takes about 40 minutes on one GPU. On the primary target, feedback is what makes the attack work: constant steering at the same average strength over the token-committing steps produces a far smaller shift while corrupting nearly three times as many outputs, and a constant strength set separately for each example still falls well short. Our findings identify the denoising trajectory as a new control channel in dLLMs and call for bias audits that examine the serving stack rather than the frozen model alone.
发表机构
- Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)
机构由 AI 辅助整理,请以论文原文为准。