发表机构
University of Turin; University of Southern Denmark(都灵大学; 南丹麦大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究探讨自然语言自编码器中如何选择最具信息量的标记位置,发现基于聊天结构的排序器优于计算信号,仅解释5%位置即可保留大部分威胁解释成功率,且预训练口头表达器可恢复隐藏单词。
AI 中文摘要
自然语言自编码器将语言模型的内部激活转换为可读的解释。解释每个标记位置成本高昂。审计员应检查哪些位置以理解潜在威胁?我们在提示注入和隐藏的470万个解释上研究此问题。我们比较了模型计算信号与仅基于聊天结构训练的排序器。聊天结构通常比计算信号选择更相关的解释,且无需为位置选择进行模型前向传递。在四个数据集中的三个上,仅解释5%的位置就保留了几乎全部的解释所有位置的成功率,其中成功意味着获得关于威胁的解释。收益因审计任务而异。我们还表明,预训练的口头表达器可以恢复模型通过微调学习隐藏的单词,而无需额外的口头表达器训练。这些结果确定了审计员可以集中生成解释的位置,并表明有用的解释可以扩展到口头表达器被训练描述的模型之外。
英文摘要
Natural language autoencoders translate a language model's internal activations into readable explanations. Explaining every token position is costly. Which positions should an auditor inspect to understand a potential threat? We study this question across $4.7$ million explanations on prompt injection and concealment. We compare signals from model computation with a ranker trained only on chat structure. Chat structure usually selects more relevant explanations than the computational signals, without requiring a model forward pass for position selection. On three of four datasets, explaining just $5\%$ of positions retains nearly all of the success rate from explaining every position, where success means obtaining an explanation about the threat. The benefit varies with the audit task. We also show that pretrained verbalizers recover words that models have learned to conceal through fine-tuning, without additional verbalizer training. These results identify where auditors can concentrate explanation generation and show that useful explanations can extend beyond the model a verbalizer was trained to describe.