arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.14791cs.AI

用于研究语言模型中欺骗行为的转码器

Transcoders for Investigating Deception in Language Models

  • Home Team Science & Technology Agency (HTX)(新加坡内政部科技局)

机构由 AI 辅助整理,请以论文原文为准。

Darius Lim, Nathan Leow, Xin Wei Chia

AI总结:

研究用转码器分析语言模型欺骗行为,通过预训练转码器的Qwen3 - 4B模型构建归因图,经特征引导和电路分析识别相关特征字典,发现欺骗源于模型内部机制,凸显转码器在监测及检测语言模型安全漏洞方面的潜力。

AI中文摘要:

转码器最近成为一种很有前景的机械可解释性(MI)方法,能对模型行为进行电路级分析。本文研究用转码器分析语言模型中的欺骗行为,该行为存在安全风险。我们使用预训练转码器的Qwen3 - 4B模型,构建捕捉特征激活和特征间依赖关系的归因图,通过特征引导和电路分析,识别出欺骗相关特征字典,发现这些特征对欺骗输出影响更强。研究表明欺骗源于模型内部机制,凸显转码器在行为监测和早期检测语言模型安全漏洞方面的潜力。

英文摘要:

Transcoders have recently emerged as a promising approach for mechanistic interpretability (MI), enabling circuit-level analysis of model behaviour. In this paper, we investigate the use of transcoders to analyse deceptive behaviour in language models, a behaviour that poses a safety and security risk. Using a Qwen3-4B model with pre-trained transcoders, specifically per-layer transcoders (PLTs), we construct attribution graphs that capture feature activations and inter-feature dependencies, allowing circuit-level analysis of deception. Through feature steering and circuit analysis, we identified a dictionary of deception-related features and show that these features exert a stronger influence on deceptive outputs, as they produce predictable shifts between deceptive and non-deceptive responses. These findings suggest that deception emerges from internal model mechanisms and highlight the potential of transcoders for behavioural monitoring and early detection of security vulnerabilities related to malicious behaviours in language models.

↑