Think Twice Before You Act: Enhancing Agent Behavioral Safety with Thought Correction
三思而后行:通过思想修正增强智能体行为安全
机构 * Fudan University, Shanghai, China(复旦大学,上海,中国) ; Shanghai Innovation Institute, Shanghai, China(上海创新研究院,上海,中国) ; Shanghai Pudong Research Institute of Cryptology, Shanghai, China(上海浦东密码研究院,上海,中国)
专题命中 规划推理 :reasoning(abstract);分类 cs.AI
AI总结 提出Thought-Aligner,一种轻量级插件式安全模型,在动作执行前对不安全思想进行因果修正,无需修改底层智能体,通过两阶段对比学习训练,在多个基准和六种LLM上将行为安全从约50%提升至约90%,超越现有防护约23%,同时提升有用性约5%。
Comments Accepted to ICML 2026