When Models Outthink Their Safety: Unveiling and Mitigating Self-Jailbreak in Large Reasoning Models
当模型超越安全限制:揭示并缓解大推理模型的自我突破
机构 * Chinese Information Processing Laboratory, Institute of Software, Chinese Academy of Sciences(中国科学院软件研究所信息处理实验室) ; University of Chinese Academy of Sciences(中国科学院大学)
AI总结 研究揭示大推理模型存在自我突破现象,提出CoG框架通过分步干预缓解安全问题,平衡安全与推理性能。
Comments ACL 2026. The first two authors contributed equally. The main text is 9 pages, with an appendix of 28 pages. The paper contains 20 figures and 15 tables