arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.35870cs.CR

说会阻止,实际仍行动:为何LLM安全判断无法约束LLM智能体的行动

Says Block, Still Acts: Why LLM Safety Judgments Fail to Govern Action in LLM Agents

Dongsheng Chen, Xiangyu Zhao, Xin Yao, Xuetao Wei

首次发表
浏览论文内容

中文总结 AI 辅助

本研究揭示LLM智能体中安全判断与行动脱节源于因果控制弱耦合,而非信息缺失,并强调安全计算需与行动选择因果耦合。

中文摘要 AI 辅助

大型语言模型智能体能够正确判断某个行动应被阻止,却仍然倾向于执行该行动。我们探究这种判断-行动脱节为何产生,以及显式安全判断是否在因果层面主导后续的行动偏好。在三个开放权重语言模型上,安全预测信息仍可从行动状态中恢复,这反驳了简单的信息丢失解释。相反,这种脱节更可由判断侧与行动侧因果控制之间的弱耦合来解释:能够可靠地将显式安全判断转向“阻止”的干预措施,对行动偏好产生的改变远小于行动原生干预措施。这种不对称性在共享的判断到行动轨迹中持续存在,其中对判断的强上游控制并未转化为对行动偏好的同等强度的下游控制。在单个干预方向之外,判断控制子空间与行动控制子空间仅部分重叠,而有效的行动控制在正交于判断控制子空间的方向上仍然可用。综合来看,这些结果区分了信息可用性与因果控制:LLM智能体可以保留识别行动不安全所需的信息,而无需让支持该判断的变量可靠地主导其行动偏好。对于智能体安全而言,这表明仅改进安全识别或自我批判可能不足,除非安全相关计算也与行动选择形成因果耦合。

英文摘要

Large language model agents can correctly judge that an action should be blocked while still preferring to take it. We ask why this judgment-action disconnect arises, and whether explicit safety judgment causally governs subsequent action preference. Across three open-weight language models, safety-predictive information remains recoverable from action states, arguing against a simple information-loss account. Instead, the disconnect is better explained by weak coupling between judgment- and action-side causal control: interventions that reliably shift explicit safety judgments toward BLOCK produce much smaller changes in action preference than action-native interventions. This asymmetry persists within a shared judgment-to-action trajectory, where strong upstream control of judgment does not translate into comparably strong downstream control of action preference. Beyond individual intervention directions, judgment- and action-control subspaces overlap only partially, while effective action control remains available in directions orthogonal to the judgment-control subspace. Together, these results distinguish information availability from causal control: an LLM agent can retain the information needed to recognize an action as unsafe without the variables supporting that judgment reliably governing its action preference. For agent safety, this suggests that improving safety recognition or self-critique alone may be insufficient unless safety-relevant computations are also causally coupled to action selection.

发表机构

  • Southern University of Science and Technology(南方科技大学)
  • City University of Hong Kong(香港城市大学)
  • Lingnan University(岭南大学)

机构由 AI 辅助整理,请以论文原文为准。

↑