AI 中文总结
针对大语言模型安全对齐的表面性问题,该研究提出潜在意图验证(LIV)方法,利用小型语言模型早期层的有害特征,在不重新训练的情况下,将语义伪装攻击的检测性能提升20%-50%。
AI 中文摘要
大语言模型(LLM)的安全对齐往往流于表面,依赖仅在生成最终阶段触发的拒绝机制,却未擦除预训练期间习得的有害概念的基础知识。本研究表明,这种架构上的脱节使模型易受“语义伪装”攻击——这类对抗性攻击将有害意图包裹在良性叙事语境(如创意写作)中,可有效绕过标准输入输出防护栏。通过分析三类不同小型语言模型(SLM)系列(Phi-3、Qwen2.5、Gemma-2b)在对抗压力下的潜在激活轨迹,本研究识别出一个通用的“意图地平线”:一个关键深度(通常为总层数的15%-20%),此时模型将查询语境化为“安全”叙事时,其预训练的有害意图的独特表征会崩溃。结果显示,伪装攻击的后期层表征与安全查询在数学上无法区分(检测率<20%),而早期层表征保留着可检测的独特“有害特征”。利用这一见解,本文提出轻量级探测防御方法“潜在意图验证(Latent Intent Verification, LIV)”。在PKU-SafeRLHF数据集上的实验表明,LIV在所有测试架构上的性能较标准防护栏提升了20%-50%,可在无需模型重新训练的情况下有效中和零日语义攻击。
英文摘要
Safety alignment in Large Language Models (LLMs) is often superficial, relying on refusal mechanisms that trigger only at the final stages of generation without erasing the foundational knowledge of harmful concepts acquired during pretraining. This study demonstrates that this architectural disconnect leaves models vulnerable to Semantic Camouflage -- adversarial attacks that wrap harmful intent in benign narrative contexts (e.g., creative writing), effectively bypassing standard input and output guardrails. By analyzing the latent activation trajectories of three distinct Small Language Model (SLM) families (Phi-3, Qwen2.5, and Gemma-2b) under adversarial stress, this research identifies a universal ``Intent Horizon'' -- a critical depth (typically 15--20\% of total layers) where the model's distinct, pre-trained representation of harmful intent collapses as it contextualizes the query into a ``safe'' narrative. Results indicate that while late-layer representations of camouflaged attacks are mathematically indistinguishable from safe queries (Detection Rate $< 20\%$), early-layer representations retain a distinct, detectable ``harm signature.'' Leveraging this insight, this paper proposes Latent Intent Verification (LIV), a lightweight probing defense. Experiments on the PKU-SafeRLHF dataset demonstrate that LIV outperforms standard guardrails by a margin of 20--50\% across all tested architectures, effectively neutralizing zero-day semantic attacks without requiring model retraining.
Comments5 pages, 3 figures. Accepted at the 2026 IEEE 2nd International Conference on Quantum Photonics, Artificial Intelligence, and Networking (QPAIN)
Journal refProc. 2026 IEEE 2nd International Conference on Quantum Photonics, Artificial Intelligence, and Networking (QPAIN), 2026
DOI:10.1109/QPAIN69676.2026.11546227