arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SeFaR:面向深度神经网络的语义特征感知鲁棒性测试

SeFaR: Semantic Feature-aware Robustness Testing of Deep Neural Networks

Nusrat Jahan Mozumder, Divya Gopinath, Corina Pasareanu, Matthew Dwyer

arXiv 2608.10289首次发表:更新:

发表机构

University of Virginia; KBR Inc.; NASA Ames; Carnegie Mellon University(弗吉尼亚大学; KBR公司; 美国国家航空航天局艾姆斯研究中心; 卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

SeFaR是一个以语义特征为核心的视觉模型鲁棒性测试框架,可在保留需求满足性的前提下,生成测试输入以识别影响模型决策的特征,进而发现故障并关联相关特征。

AI 中文摘要

深度神经网络越来越多地被部署在安全关键领域作为感知模块,其故障常由罕见及代表性不足的场景导致,因此需要评估感知模型的语义鲁棒性,即模型行为在真实感知变异下对高层需求的符合程度。为解决该问题,本文提出SeFaR,一个以语义特征为核心的视觉模型系统性测试框架。给定自然语言需求及一组满足需求的输入,SeFaR会在保留需求满足性的前提下,针对多样化的真实语义变异评估鲁棒性。该方法采用新颖的分层概念模型,可对特征空间进行结构化探索,并通过用户定义的概念融入领域知识;利用当前最先进的扩散模型及视觉-语言模型生成保语义的逼真扰动,同时识别影响模型行为的未知特征;采用反馈驱动的自适应流程生成可解释的、诱发故障的语义概念及对应测试输入。案例研究评估表明,该框架可有效满足需求前提,同时识别出影响模型决策的与需求无关的特征,既能发现故障,又能将故障与这类特征关联起来。

英文摘要

Deep neural networks are increasingly deployed in safety-critical domains as perception modules, where failures are often caused due to rare and under-represented scenarios. This necessitates the need to evaluate the semantic robustness of perception models; conformance of behavior to high-level requirements over real-world perceptual variability. To address this, we propose SeFaR, a framework for systematic semantic-feature-centric testing of vision models. Given a natural-language requirement and a set of satisfying inputs, SeFaR evaluates robustness with respect to diverse realistic semantic variations that preserve requirement satisfaction. The approach employs a novel hierarchical concept model enabling structured exploration of the feature space and incorporation of domain knowledge via user-defined concepts. State-of-the-art diffusion and vision-language models are leveraged to generate photorealistic semantics-preserving perturbations and identification of previously unknown features impacting behavior. A feedback-driven adaptive process is adopted to generate interpretable failure-inducing semantic concepts along with corresponding test inputs. Evaluation on case studies demonstrates that the proposed framework effectively satisfies requirement preconditions while identifying requirement-independent features that influence model decisions, enabling it to both uncover faults and relate them to such features.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑