发表机构
Hong Kong Baptist University; Guangdong Polytechnic Normal University; Guangdong Institute of Digital Industry; The Hong Kong Polytechnic University; The Education University of Hong Kong; Southern University of Science and Technology(香港浸会大学; 广东技术师范大学; 广东数字产业研究院; 香港理工大学; 香港教育大学; 南方科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
HoloAegis是一种最小参数拓扑推理框架,通过冻结语义表示的纯几何推理实现零样本LLM安全护栏,在8个基准测试中达到最先进准确率,兼具低延迟、零冷启动数据和跨语言迁移能力。
AI 中文摘要
当前大语言模型(LLM)安全护栏面临一个根本矛盾:微调会扭曲预训练表示,而生成式评判器则会产生过高的推理成本。我们挑战这一主流范式,提出问题:能否仅通过对冻结语义表示的纯几何推理来实现安全?我们提出HoloAegis,这是一种最小参数拓扑推理框架,它将表示与推理解耦。我们称我们的方法为最小参数,因为唯一的自由参数是锚点数量K和温度tau,两者在构建后固定,无需基于梯度的训练。未经过微调的编码器将文本映射到单位球面上,之后所有决策均为纯几何操作。我们将安全评估形式化为对预计算系统拓扑锚定库的吉布斯-玻尔兹曼自由能计算,并引入双时间尺度指数移动平均来检测渐进式多轮语义漂移。我们的关键理论见解是拓扑边界稳定性猜想:我们提供理论依据和有力的经验证据,表明稀疏锚点质心比全向量空间方法能更好地稳定决策边界,抵御高频词汇扰动。在8个基准测试中,HoloAegis达到了最先进的准确率(AuthenHallu上的AUC为1.0000,HarmBench上为0.9802),具有亚毫秒级延迟、零冷启动数据以及跨语言迁移能力(中文CHIFRAUD上的AUC为0.9758)。
英文摘要
Current LLM safety guardrails face a fundamental tension: fine-tuning distorts pre-trained representations while generative judges incur prohibitive inference costs. We ask a complementary question: how far can safety be achieved through pure geometric reasoning over frozen representations, and where does it fail? We present HoloAegis, a minimally parametric topological inference framework that decouples representation from reasoning: an un-fine-tuned encoder maps text to the unit sphere S^{d-1}, and all decisions reduce to Gibbs-Boltzmann free-energy differences over pre-computed anchor centroids. We contribute a boundary-mapping study rather than a leaderboard claim. On a frozen three-benchmark protocol, HoloAegis (3.2 MB) statistically matches WildGuard-7B (14 GB) on toxicity (0.96 vs. 0.96), exceeds it on harmful behaviors (0.99 vs. 0.79), and cedes oversafety detection (0.62 vs. 0.98) -- while ShieldGemma-2B fails on indirect harms (0.34). These failure modes are complementary and mechanistically traceable: potential-difference scoring senses manifold clustering, whereas policy-conditioned LLM judging requires explicit taxonomy matching. We restate our Topological Boundary Stability conjecture in ratio form and validate it via reference-set bootstrap: anchor banks reduce score variance 4-15x and boundary displacement to approximately 0.44 + 0.23 sqrt(k/K) of the full-space estimator. Per-domain analysis further reveals that geometric separability tracks within-domain semantic homogeneity. Our results chart where geometric guardrails substitute for, and where they must defer to, LLM judges.
CommentsPreprint v2, September 2026. 4 figures, 12 tables. Corrected and substantially revised from v1 (arXiv:2608.08485v1)