发表机构
Faculty of Electrical and Computer Engineering, Technion(以色列理工学院电气与计算机工程学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对LLMs在不道德场景下的合规性失败问题,提出LRP引导的解码方法,发现模型对提示词元的归因不足是关键原因,干预后可提升响应安全性。
AI 中文摘要
尽管大语言模型(LLMs)经过对齐以优化有用性和无害性这两个目标,但这两个目标可能会发生冲突,不可避免地导致对齐失败。本研究系统调查了LLMs未能表现出道德行为的实例。为了理解这些漏洞的潜在机制,我们引入了一种探测方法,以三种不同的结构模态向LLMs呈现不道德场景:客观分类任务、主观第一人称陈述以及直接的协助请求。我们发现,基于协助请求的形式下模型性能会下降。利用逐层相关性传播(LRP),我们将这种差异追溯到归因偏差:模型更强调良性的任务框架词元(例如“你能帮我……吗”),而非表明潜在不道德行为的词元(例如“不被抓住”),我们将后者称为提示词元。我们假设这种归因不足会导致有害的合规性。为了验证这一点,我们引入了两种LRP引导的解码方法,引导生成过程向与提示词元更相关的轨迹发展。实证评估表明,这些干预措施能促进更安全的响应,支持提示词元归因在合规失败中的作用。
英文摘要
Although Large Language Models (LLMs) are aligned to optimize for both helpfulness and harmlessness, these dual objectives may conflict, inevitably leading to alignment failures. This work systematically investigates instances where LLMs fail to exhibit ethical behavior. To understand the underlying mechanics of these vulnerabilities, we introduce a probing methodology that presents unethical scenarios to LLMs in three distinct structural modalities: objective classification tasks, subjective first-person statements, and direct requests for assistance. We find that model performance degrades in the request-for-assistance-based form. Using Layer-wise Relevance Propagation (LRP), we trace this discrepancy to an attribution bias: the model places greater emphasis on benign task-framing tokens (e.g., "Can you help me...") than on tokens signaling the underlying unethical behavior (e.g., "without getting caught"), which we term cue-tokens. We hypothesize that this under-attribution contributes to harmful compliance. To test this, we introduce two LRP-guided decoding methods that steer generation toward trajectories more relevant to cue tokens. Empirical evaluations show that these interventions promote safer responses, supporting cue-token attribution's role in compliance failures.
CommentsSocialAgent, NeurIPS 2026
Journal refNeurIPS 2026, SocialAgent Workshop