Before the Last Token: Diagnosing Final-Token Safety Probe Failures
在最后的标记之前:诊断最终标记安全探针失败
机构 * SafeSwitch ; HarmBench ; SorryBench
AI总结 研究最终标记安全探针在预填充阶段的失败模式,发现安全证据分布在早期用户标记中,传统探针无法捕捉,提出基于PCA-HMM的轨迹模型提升诊断能力。
Comments 8 pages, 2 figures, 7 tables. Accepted at the ICML 2026 Mechanistic Interpretability Workshop and the ICML 2026 Failure Modes in Agentic AI Workshop