发表机构
KU Leuven(鲁汶大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究揭示工具中介交互使LLM拒绝有害请求的阈值更高且更脆弱,表明工具环境可能降低模型安全性,常规评估需调整。
AI 中文摘要
大型语言模型(LLMs)越来越多地被部署为可访问外部工具,然而与常规对话交互相比,有害的工具中介交互被拒绝的可能性更低。由于这种拒绝行为的变化仍未得到充分探索,我们在一组多样化的开放权重语言模型中调查了其潜在机制。我们发现,关于请求有害性的信息仍然强烈地编码在模型的表示中,并在对话和工具中介输入之间转移。表示几何和神经元层面的分析证据进一步表明,这两种交互模式系统性地以不同方式分配与危害相关的计算。关键的是,虽然对话输入可以在相对较低的有害性感知水平上被拒绝,但工具中介输入在有害性跨越一个显著更高的有效拒绝阈值之前仍保持允许状态。此外,工具中介的拒绝也更为脆弱:逐步削弱拒绝计算会在比对话拒绝更低的干预强度下破坏工具中介的拒绝,即使良性能力保持完整。总之,我们的发现表明,工具中介并不仅仅降低对危害的内部感知,而是影响其转化为拒绝。总体而言,这表明工具中介环境可能从根本上降低模型对有害请求的鲁棒性,并且常规的安全评估可能无法完全适用于LLM智能体。
英文摘要
Large language models (LLMs) are increasingly deployed with access to external tools, yet harmful tool-mediated interactions are less likely to be refused when compared to regular conversational ones. As this change in refusal behavior remains underexplored, we investigate its underlying mechanisms across a diverse set of open-weight language models. We find that information about the harmfulness of a request remains strongly encoded in the model's representations and transfers across conversational and tool-mediated inputs. Evidence from representation geometry and neuron-level analysis further indicates that the two interaction modes systematically distribute harm-related computation differently. Crucially, while conversational inputs can be refused at relatively low levels of perceived harmfulness, tool-mediated inputs remain permissive until harmfulness crosses a substantially higher effective refusal threshold. Moreover, tool-mediated refusal is also more brittle: progressively weakening the refusal computation disrupts tool-mediated refusal at lower intervention strengths than conversational refusal, even when benign capabilities remain intact. Together, our findings indicate that tool mediation does not simply reduce the internal perception of harm, but instead impacts its conversion into refusal. Overall, this suggests tool-mediated environments may intrinsically reduce robustness of models to harmful requests, and that conventional safety evaluations may not fully transfer to LLM agents.