arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TestifAI:基于层析成像的深度学习系统测试方法

TestifAI: Tomography-Based Testing for Deep Learning Systems

Arooj Arif, Tobias Hartung, Elena Botoeva, Alexandros Koliousis

arXiv 2608.18900首次发表:更新:

发表机构

Northeastern University London; University of Kent(伦敦东北大学; 肯特大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

TestifAI是一种深度学习测试框架,通过部分模型层析成像从低阶扰动观测预测高阶扰动测试结果,可高效准确估计多扰动下的模型鲁棒性,减少60-80%的推理量且总误差低于7%。

AI 中文摘要

随着AI系统越来越多地部署在自动驾驶等安全关键应用领域,相关风险也随之增加。因此,支撑现代AI系统的深度学习模型必须经过全面测试以确保其正确运行。单次鲁棒性测试需要进行数千次推理,以实证验证模型在输入有界扰动下的输出是否保持稳定。然而,现有的测试框架缺乏系统探索和总结组合扰动空间下鲁棒性的手段。我们提出TestifAI,这是一种深度学习测试框架,用于高效且准确地估计模型对组合扰动的鲁棒性。TestifAI允许用户将操作条件指定为语义输入扰动的结构化空间(例如图像模糊、亮度和缩放)以及离散的严重程度级别(例如低、中、高),用户可查询任意组合(例如“低模糊、高亮度和中缩放”)的模型鲁棒性。为实现高效与准确,TestifAI引入了部分模型层析成像,这是一种仅通过应用少量扰动(低阶投影)的测试,在多扰动空间中重建模型行为的新方法。为估计对至少三种扰动的鲁棒性,TestifAI仅基于涉及最多两种扰动的测试结果训练辅助模型,避免执行指数级数量的测试。我们在五个图像和语言分类任务上的实验表明,TestifAI可从低阶(1和2种扰动)观测预测高阶(3和4种扰动)测试结果,鲁棒性估计的总误差低于7%,同时将推理数量减少60-80%。

英文摘要

As AI systems are increasingly deployed in safety-critical application domains (e.g., autonomous driving), associated risks increase too. Deep learning models underlying modern AI systems, therefore, must undergo thorough testing to ensure their correct behaviour. A single robustness test involves thousands of inferences to empirically verify if a model's outputs remain stable under a bounded perturbation of its inputs. However, existing testing frameworks lack the means to systematically explore and summarise robustness across a combinatorial space of perturbations. We propose TestifAI, a deep learning testing framework for efficient and accurate estimation of robustness against combinations of perturbations. TestifAI enables users to specify operational conditions as structured spaces of semantic input perturbations (e.g., image blur, brightness and zoom) and discrete severity levels (e.g., low, medium and high). Users can query model robustness for any combination (e.g., "low blur, high brightness, and medium zoom"). To achieve efficiency and accuracy, TestifAI introduces partial model tomography, a novel approach to reconstructing model behaviour in a multi-perturbation space from tests that apply only a small number of perturbations (lower-order projections). To estimate robustness against at least three perturbations, TestifAI trains an auxiliary model on the results of tests involving up to two perturbations only, avoiding execution of an exponential number of tests. Our experiments on five image and language classification tasks show that TestifAI can predict higher-order (3 and 4 perturbations) test outcomes from low-order (1 and 2 perturbations) observations with an aggregate robustness estimation error of less than 7%, while reducing the number of inferences by 60-80%.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑