发表机构
Wrynx Inc; INTI International College Penang(Wrynx公司; 槟城英迪国际学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究复现了轻量MLP安全探针的性能,验证其可泛化至多类模型家族,且发现推理种子不影响最终标记隐向量,为安全探针的可复现性提供了新结论。
AI 中文摘要
Khatri等人(2026)[DOI: https://doi.org/10.1109/DSN-W70714.2026.00027]的研究表明,针对单个8B规模模型LLaMA-3.1-8B的最终层激活训练轻量多层感知机(MLP)探针,可检测有害提示,其F1值与规模大1000倍的安全防护模型相当,且每个基准仅用一个探针。我们复现了该流程的端到端实现,并从原研究未涉及的两个维度进行扩展:其一,通过在Gemma-4-E4B、Mistral-7B-v0.3、Qwen2-7B等模型的激活上训练相同探针,在WildJailbreak、BeaverTails、AEGIS 2.0三个基准上测试该结果是否可泛化至其他模型架构与规模;其二,通过在5个随机种子下重复激活提取并测量F1分数的方差,测试报告的性能受推理阶段非确定性的影响程度。我们的结果复现了原LLaMA模型基准的F1分数,与原结果的偏差在0.37个百分点以内(BeaverTails基准偏差在0.2个百分点以内)。研究发现,原MLP探针架构可扩展至其他模型家族,其F1分数与LLaMA-3.1-8B的报告值偏差在1个百分点以内。种子值变化的实验揭示了一个有趣现象:无论使用何种种子,所有测试架构的最终标记隐向量均保持一致。
英文摘要
Khatri et al. (2026) [DOI: 10.1109/DSN-W70714.2026.00027] show that lightweight MLP probes on final-layer activations of a single 8B model (LLaMA-3.1-8B) detect harmful prompts at F1 competitive with guard models 1000x larger, using one probe per benchmark. We reproduce this pipeline end-to-end and extend it along two axes the original study leaves open. First, we test whether the result generalizes across other model architecture and scale by training identical probes on activations from models like Gemma-4-E4B, Mistral-7B-v0.3, and Qwen2-7B, using the three benchmarks (WildJailbreak, BeaverTails, AEGIS 2.0). Second, we test how much of the reported performance is affected by non-determinism during inference by repeating extraction under five random seeds and measuring the variance of F1 scores. Our results reproduce the original LLaMA model benchmarks within 0.37 percentage points of the original F1 scores (and within 0.2 points on BeaverTails). We find that the original MLP probe architecture extends to other model families with F1 scores within a point of the values reported for LLaMA-3.1-8B. Our experiments varying seed values reveal an interesting observation: final token latent vectors remained the same for all tested architectures irrespective of the seed values used.