Training-Free Policy Violation Detection via Activation-Space Whitening in LLMs
无需训练的策略违规检测:通过激活空间白化在大语言模型中
机构 * Fujitsu Research of Europe(富士通欧洲研究机构) ; Ben-Gurion University of the Negev(贝内杰尔大学)
AI总结 本文提出一种无需训练的策略违规检测方法,通过激活空间白化技术在大语言模型中实现高效检测。
Comments Accepted to the AAAI 2026 Deployable AI (DAI) Workshop