arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

输入层饥饿:为什么逐层剪枝会破坏物联网入侵检测器

Input-Layer Starvation: Why Per-Layer Pruning Breaks IoT Intrusion Detectors

Md Anas Biswas

arXiv 2609.30729首次发表:更新:

发表机构

University of Portsmouth(朴茨茅斯大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究发现逐层剪枝导致物联网入侵检测器输入层饥饿,引发类别级失效和误报,保护首层权重或全局剪枝可预防,重算归一化统计量可修复。

AI 中文摘要

针对小型物联网(IoT)设备的入侵检测器通常通过剪枝进行压缩,并以总体准确率作为评判标准。我们表明,这种做法掩盖了一种严重的类别级失效,我们找到了其原因,并提供了低开销的预防和修复方法。在CICIoT2023数据集上,一个采用均匀逐层幅度剪枝、稀疏度为80%的两层卷积检测器,其准确率下降了16个百分点,但宏F1分数(各类别F1的平均值)下降了一半(在五个独立训练的模型上从0.542降至0.271);34个类别中有17个类别受到实质性损害。剩余的权重数量并不能解释这一现象:一个感知机和一个Transformer在剪枝到相同或更少权重时,损失最多为0.096。问题出在第一层。该层有192个权重;均匀剪枝后仅剩38个,其64个滤波器中有46%失去了所有输入权重,在这种饥饿状态下进行微调,会使第一个归一化层的运行均值在少数幸存的通道中偏移多达0.8个标准差,而部署的模型正是在这些通道上崩溃。保护这192个权重,或在相同稀疏度下进行全局剪枝,可以防止崩溃(损失为0.013);在未标记的训练数据上重新计算归一化统计量,且不改变任何权重,可以修复该问题(损失为0.039),并将误报率恢复到33%(密集模型为29%)。损害在第一层稀疏度上表现出强烈的递增剂量-反应关系,使感知机的输入层饥饿也能重现这种崩溃,且该模式在TON_IoT数据集上同样成立。这种失效表现为错误归因和误报,而非静默逃逸:在验证集选定的盲区上,均匀剪枝的检测器对72%的流量进行了错误归因,而第一层受保护的模型为50%,密集模型为47%。

英文摘要

Intrusion detectors for small Internet-of-Things (IoT) devices are usually compressed by pruning and judged by overall accuracy. We show that this hides a severe class-level failure, find its cause, and give low-overhead prevention and repair. On CICIoT2023, a two-layer convolutional detector pruned with uniform layer-wise magnitude pruning at 80% sparsity loses 16 points of accuracy but half of its macro-F1, the mean per-class F1 (0.542 to 0.271 over five independently trained models); 17 of 34 classes are materially damaged. Remaining weight count does not explain it: a perceptron and a transformer pruned to the same or fewer weights lose at most 0.096. The first layer does. It has 192 weights; uniform pruning leaves 38, 46% of its 64 filters lose every input weight, and fine-tuning under that starvation leaves the running means of the first normalisation layer displaced by up to 0.8 standard deviations in a few surviving channels, on which the deployed model collapses. Protecting those 192 weights, or pruning globally at the same sparsity, prevents the collapse (loss 0.013); recomputing the normalisation statistics on unlabelled training data, with no weight changed, repairs it (loss 0.039) and returns the false-alert rate to 33% (dense 29%). Damage shows a strong increasing dose-response in first-layer sparsity, starving a perceptron's input layer reproduces the collapse, and the pattern holds on TON_IoT. The failure is misattribution and false alerts, not silent evasion: on validation-selected blind spots, uniformly pruned detectors misattribute 72% of the traffic, against 50% with the first layer protected and 47% for the dense model.

Comments38 pages including a 15-page supplement; 12 main tables, 4 figures. Code and results: https://github.com/anasbiswas1/iot-trust-compression

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑