arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过神经崩溃不稳定性对深度神经网络进行测试用例优先级排序

Test Case Prioritization for DNNs via Neural Collapse Instability

Chunyu Liu, Mingyuan Li, Yang Li, Wenmin Li, Fei Gao, Tengfei Tu, Su-Juan Qin

arXiv 2607.20046首次发表:更新:

发表机构

Beijing University of Posts and Telecommunications(北京邮电大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对深度神经网络在安全关键领域部署时,有限测试预算下模型验证成本高的问题,提出 NCIP 框架,利用跨检查点预测变异性取代绝对置信度,通过特定组件选择检查点及对测试输入排序,实验表明该框架在早期故障发现方面性能优于基线。

AI 中文摘要

随着深度神经网络(DNN)在安全关键领域的广泛部署,在有限测试预算下降低模型验证成本变得愈发重要。现有测试用例优先级排序技术常依赖输出概率的单检查点置信信号,但DNN可能自信地犯错且预测类与竞争类之间的置信 margin 小,削弱早期故障发现。我们提出神经崩溃启发的优先级排序(NCIP)框架,在终端训练阶段用跨检查点预测变异性取代绝对置信度,此阶段模型几何结构高度结构化。NCIP 引入两个关键组件,一是用分类器权重的等角分数选择训练检查点的 NC 引导代表性子集,二是按跨所选检查点的预测变异性对测试输入排序,找出在检查点引起的决策边界移动下不稳定的边界相邻和易失败样本。多个数据集和架构的大量实验表明,与竞争基线相比,NCIP 在早期故障发现方面性能强劲,在相同测试预算下,RAUC - ALL 增益为 1.5%至 16.6%,RAUC - 500 增益为 4.9%至 20.6%,且在所有数据集 - 模型对中平均性能最佳。

英文摘要

With the widespread deployment of deep neural networks (DNNs) in safety-critical domains, reducing the cost of model validation under limited testing budgets has become increasingly important. Existing test case prioritization techniques often rely on single-checkpoint confidence signals derived from output probabilities. However, DNNs can be confidently wrong, and the confidence margin between the predicted and competing classes is frequently small, which weakens early fault discovery. To address this limitation, we propose a Neural-Collapse-Inspired Prioritization (NCIP) framework that replaces absolute confidence with cross-checkpoint prediction variability in the terminal training regime, where model geometry becomes highly structured. NCIP introduces two key components. First, it selects an NC-guided representative subset of training checkpoints using an equiangularity score of classifier weights, quantified as the standard deviation of pairwise cosine similarities among class weight vectors. Second, it prioritizes test inputs by their prediction variability across the selected checkpoints, surfacing boundary-adjacent and failure-prone samples that are unstable under checkpoint-induced decision boundary shifts. Extensive experiments across multiple datasets and architectures show that NCIP achieves strong performance in early fault discovery compared with competitive baselines, with 1.5 to 16.6 percent RAUC-ALL gains and 4.9 to 20.6 percent RAUC-500 gains under the same testing budget. NCIP further attains the best average performance across all dataset-model pairs.

Comments24 pages. Accepted at ISSTA 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑