arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TESTNAV:面向组合鲁棒性测试的帕累托引导搜索

TESTNAV: Pareto-Guided Search for Compositional Robustness Testing

Arooj Arif, Tobias Hartung, Elena Botoeva, Alexandros Koliousis

arXiv 2608.19882首次发表:更新:

发表机构

Northeastern University London; University of Kent(伦敦东北大学; 肯特大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

TESTNAV是帕累托引导的鲁棒性测试框架,采用NSGA-II近似双目标帕累托前沿,在四类跨模态基准上,以35.8%-89.3%的扰动空间实现比基线快2.15倍的帕累托前沿恢复。

AI 中文摘要

深度学习模型仍易受现实世界输入扰动影响,尤其是同一输入中同时存在多种损坏(如亮度偏移与运动模糊)时。组合测试可揭示这些交互效应,但带来两大挑战:随扰动维度与严重程度等级增加,扰动空间呈组合式增长;且诊断价值不均——许多组合生成的输入退化不切实际,实用相关性有限。我们提出TESTNAV,一种帕累托引导的鲁棒性测试框架,用于在仅能评估有限数量扰动配置时,高效探索离散组合扰动空间。TESTNAV通过将鲁棒性测试建模为双目标优化,优先考虑严重且真实的故障:最大化性能退化,同时保留由模态特定指标(如视觉任务的SSIM与KID、语言与代码任务的chrF与BERT-F1)衡量的输入保真度。它采用NSGA-II算法近似双目标帕累托前沿。在涵盖视觉、自然语言与代码生成的四个基准测试中,TESTNAV恢复帕累托前沿的速度比基于搜索的基线快达2.15倍,仅使用由四个扰动维度(每个维度含六个等级)定义的离散扰动空间的35.8%-89.3%。

英文摘要

Deep learning models remain vulnerable to real-world input perturbations, especially when multiple corruptions co-occur in the same input (e.g., brightness shifts and motion blur). Compositional testing reveals these interaction effects but introduces two challenges: combinatorial growth of the perturbation space as dimensions and severity levels increase, and uneven diagnostic value-many combinations yield unrealistically degraded inputs with limited practical relevance. We present TESTNAV, 1 a Pareto-guided robustness testing framework for efficiently exploring discrete, compositional perturbation spaces when only a limited number of perturbation configurations can be evaluated. TESTNAV prioritises severe yet realistic failures by formulating robustness testing as bi-objective optimisation: maximise performance degradation while preserving input fidelity measured by modality-specific metrics (e.g., SSIM and KID for vision; chrF and BERT-F1 for language and code). It uses NSGA-II to approximate the bi-objective Pareto front. Across four benchmarks spanning vision, natural language, and code generation, TESTNAV recovers Pareto fronts up to 2.15x faster than search-based baselines, using 35.8%-89.3% of the discrete perturbation space defined by four perturbation dimensions with six levels each.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑