arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于XAI引导搜索的DNN机器人导航系统泛化鲁棒性测试

Generalizable Robustness Testing of DNN-Based Robotic Navigation Systems via XAI-Guided Search

Khizra Sohail, Miren Illarramendi, Aitor Arrieta

arXiv 2610.06862首次发表:更新:

发表机构

Mondragon University(蒙德拉贡大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出XAI引导的多目标进化搜索,结合多图像优化与积分梯度,提升DNN机器人导航系统鲁棒性测试的泛化性,成功率从53.85%提升至70%,并验证仿真到物理迁移的有效性。

AI 中文摘要

**背景:**深度神经网络(DNN)日益控制网络物理系统(CPS),但微小的输入扰动可能导致不安全的系统级行为。现有方法通常针对单个图像优化扰动,且仅在仿真中评估,限制了其泛化性和实际有效性。**目标:**本研究旨在生成在操作观测中仍保持有效的鲁棒性测试,并评估所产生的失败是否从仿真迁移到物理机器人。**方法:**我们提出一种可解释性引导的多目标进化方法,通过视觉和行为聚类选择代表性图像,生成稀疏扰动。聚合积分梯度引导变异朝向有影响力的图像区域。我们在Gazebo中对DNN控制的LeoRover进行评估,开展消融研究,并在物理机器人上验证分层扰动的子集。**结果:**该方法的中位成功率达到了70.0%,而无引导搜索为53.85%,中位超体积从0.65提升至0.73。多图像优化将成功率从50.0%提升至57.5%,而XAI引导进一步将其提升至70.0%。在仿真到现实评估中,仿真达到了0.95的精确率和0.67的召回率,仿真与物理失败时间显示出显著的正相关(0.617)。**结论:**将多图像优化与可解释性引导搜索相结合,可改进DNN控制机器人系统的鲁棒性测试。仿真能有效识别并优先排序可迁移的失败,但物理验证仍然必要,因为某些现实世界失败无法在仿真中复现。

英文摘要

**Context:** Deep Neural Networks (DNNs) increasingly control Cyber-Physical Systems (CPSs), yet small input perturbations can cause unsafe system-level behavior. Existing approaches often optimize perturbations for individual images and evaluate them only in simulation, limiting their generalizability and practical validity. **Objectives:** This work aims to generate robustness tests that remain effective across operational observations and to evaluate whether the resulting failures transfer from simulation to a physical robot. **Methods:** We propose an explainability-guided multi-objective evolutionary approach that generates sparse perturbations over representative images selected through visual and behavioral clustering. Aggregated Integrated Gradients guide mutations toward influential image regions. We evaluate the approach on a DNN-controlled LeoRover in Gazebo, conduct an ablation study, and validate a stratified subset of perturbations on the physical robot. **Results:** The approach achieved a median success rate of 70.0%, compared with 53.85% for unguided search, and increased median hypervolume from 0.65 to 0.73. Multi-image optimization improved the success rate from 50.0% to 57.5%, while XAI guidance further increased it to 70.0%. In the sim-to-real evaluation, simulation achieved 0.95 precision and 0.67 recall, and simulated and physical failure times showed a significant positive correlation of 0.617. **Conclusion:** Combining multi-image optimization with explainability-guided search improves robustness testing for DNN-controlled robotic systems. Simulation effectively identifies and prioritizes transferable failures, but physical validation remains necessary because some real-world failures are not reproduced in simulation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑