arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Z-PEFT:基于规范谱特征的参数高效微调模型零样本后门检测

Z-PEFT: Zero-shot Backdoor Detection in Parameter-Efficient Fine-Tuning via Canonical Spectral Signatures

Nicola Pitzalis, Donald Shenaj, Giacomo Cignoni, Andrea Cossu, Davide Bacciu, Antonio Carta

arXiv 2608.02271首次发表:更新:

AI 中文总结

Z-PEFT是一种仅依赖逐层谱测度的轻量元分类器,可在含未见攻击和数据集的新条件下实现最优权重空间后门检测,且计算成本低、可扩展。

AI 中文摘要

参数高效微调(PEFT)模型常被从业者从开放仓库下载,这种广泛做法形成了显著攻击面,攻击者可发布带后门的模型,使其对预定义触发词产生特定行为。我们研究权重空间后门检测问题,即检测器分类器仅用模型权重预测其是否为恶意,以实现轻量安全机制。现有多数方法在封闭世界设置中设计与评估,检测器在相同攻击类型上训练和测试;相比之下,我们在包含未见攻击和数据集的新条件下评估后门检测。我们提出Z-PEFT,一种仅依赖逐层谱测度分类的轻量元分类器。实验表明,封闭世界设置下的强性能未必能迁移到零样本后门检测中;在权重空间检测器中,Z-PEFT实现了最佳性能,同时保持低且可扩展的计算成本。

英文摘要

Parameter-Efficient Fine-tuned (PEFT) models are frequently downloaded from open repositories by practitioners. This widespread practice creates a significant attack surface, as malicious actors can publish backdoored models that induce specific behaviors in response to predefined triggers. We study the problem of weight-space backdoor detection, where a detector classifier predicts whether a model is malicious using only its weights, enabling a lightweight safety mechanism. Most existing methods are designed and evaluated in a closed-world setting, where the detector is trained and tested on the same attack type. In contrast, we evaluate backdoor detection under novel conditions, including previously unseen attacks and datasets. We propose Z-PEFT, a lightweight meta-classifier that relies exclusively on layer-wise spectral measures for classification. Our experiments show that strong performance in the closed-world setting does not necessarily translate to high accuracy in zero-shot backdoor detection. Among weight-space detectors, Z-PEFT achieves the best performance while maintaining low and scalable computational cost.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑