arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34703cs.SE

当模糊性遇到非典型性:DNNs 的双视角测试输入优先级排序

When Ambiguity Meets Atypicality: Dual-Perspective Test Input Prioritization for DNNs

发表机构北京航空航天大学 · 浙江大学
查看机构详情
  • Beihang University(北京航空航天大学)
  • Zhejiang University(浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

Haoran Li, Shihai Wang, Bin Liu, Jialuo Chen, Wenjing Zhu, Yu Liu, Tengfei Shi, Shudi Guo

首次发表
浏览论文内容

中文总结 AI 辅助

针对DNN测试输入优先级排序中单一视角的盲区,提出基于KNN密度的DuFP方法,融合类间模糊性与类内非典型性,实验证明其有效优于现有方法。

中文摘要 AI 辅助

尽管深度神经网络(DNNs)在前沿领域取得了显著进展,但其固有的脆弱性已成为日益受到关注的问题。为确保基于 DNN 的软件可靠性和安全性,DNN 测试已成为不可或缺的实践。在此背景下,测试输入优先级排序对于早期故障检测和降低标注成本至关重要。然而,准确识别导致故障的输入仍然具有挑战性。尽管决策模糊性和分布非典型性是两种广泛采用的观点,分别用于刻画类间竞争和类内典型性,但仅依赖任一观点不可避免地会引入盲区。本文提出了 DuFP(双视角特征空间优先级排序),一种基于 KNN 密度的 DNN 测试输入优先级排序方法,该方法联合考虑了类间和类内两种视角。DuFP 的优先级排序框架建立在类条件密度估计之上。基于估计结果,预测正确性通过模糊性分数和非典型性分数来表征,前者反映决策模糊性,后者量化分布非典型性。随后,通过整合这两个分数构建一个混合不确定性分数,以指导最终的优先级排序。我们在干净、损坏和对抗性场景下,对图像和文本数据集上的优先级排序和选择任务评估了 DuFP。实验结果表明,DuFP 能有效且高效地对导致故障的输入进行优先级排序,并优于现有最先进的方法。

英文摘要

While Deep Neural Networks (DNNs) have achieved remarkable progress in cutting-edge domains, their inherent brittleness has become a growing concern. To ensure the reliability and safety of DNN-enabled software, DNN testing has emerged as an indispensable practice. Within this context, test input prioritization is essential for early fault detection and reducing labeling costs. However, it remains challenging to accurately identify failure-inducing inputs. Although decision ambiguity and distributional atypicality are two widely adopted perspectives for characterizing inter-class competition and intra-class typicality respectively, relying on either perspective in isolation inevitably introduces blind spots. In this paper, we propose DuFP (Dual perspective Feature space Prioritization), a KNN density-based test input prioritization approach for DNNs that jointly incorporates both inter-class and intra-class perspectives. The prioritization framework of DuFP is built upon class-conditional density estimation. Based on the estimation results, prediction correctness is characterized by an ambiguity score and an atypicality score, with the former reflecting decision ambiguity and the latter quantifying distributional atypicality. A hybrid uncertainty score is then constructed by integrating both scores to guide the final prioritization. We evaluate DuFP on prioritization and selection tasks across image and text datasets under clean, corrupted, and adversarial scenarios. Experimental results demonstrate that DuFP effectively and efficiently prioritizes fault-inducing inputs and outperforms state-of-the-art approaches.

补充信息

↑