arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.27348cs.CL

意图作为工具:轻松追踪智能体失配问题

INTENT-AS-A-TOOL Makes it Easy to Track Agentic Misalignment

  • Tsinghua University(清华大学)
  • MatrixOrigin(矩阵起源)
  • SiliconProspect AI(硅景人工智能)
  • Nanyang Technological University(南洋理工大学)

机构由 AI 辅助整理,请以论文原文为准。

Yutong Zhang, Jianshuo Dong, Peng Xu, Long Wang, Jie Zhang, Tianwei Zhang, Xiaoping Zhang, Han Qiu

AI总结:

本研究针对智能体失配问题,提出INTENT-AS-A-TOOL方法,通过添加意图专用工具追踪智能体推理过程中的意图变化,补充CoT监控,为智能体安全追踪提供细粒度信号。

AI中文摘要:

随着大语言模型(LLMs)被部署为自主智能体,安全故障越来越涉及具有重要后果的行动。我们研究智能体失配问题,即智能体在目标冲突和压力下采取有害行动。通过思维链(CoT)监控,我们发现有害执行往往在推理过程中存在意图信号。然而,事后CoT标签过于粗糙,无法显示生成过程中意图的变化。我们提出INTENT-AS-A-TOOL,这是一种添加了针对意图的工具的方法,为模型提供了表达对目标行为承诺的专用通道。调用意图工具的概率提供了无评判者、细粒度的信号,反映模型追求该行为的倾向。我们的结果表明,INTENT-AS-A-TOOL可补充CoT监控,将事后CoT标签扩展为密集轨迹,并识别在线干预的关键步骤。这些发现表明,行动偏好对于在推理过程中追踪智能体失配问题是有用的。我们的代码和数据可访问:this https URL。

英文摘要:

As large language models (LLMs) are deployed as autonomous agents, safety failures increasingly involve consequential actions. We study agentic misalignment, where agents take harmful actions under goal conflicts and pressures. Using chain-of-thought (CoT) monitoring, we find that harmful execution is often preceded by intent signals in reasoning. However, post-hoc CoT labels are too coarse to show how intent changes during generation. We introduce INTENT-AS-A-TOOL, an approach that adds intent-targeted tools to give the model a dedicated channel for expressing commitment to a target behavior. The probability of calling an intent tool provides a judge-free, fine-grained signal of the model's tendency to pursue that behavior. Our results show that INTENT-AS-A-TOOL complements CoT monitoring, expands post-hoc CoT labels into dense trajectories, and identifies critical steps for online intervention. These findings suggest that action preferences are useful for tracking agentic misalignment during reasoning. Our code and data are accessible: https://github.com/RebeccaZhang22/intent-as-a-tool.

↑