arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.11870cs.RO

通过显著性引导增强提升行为克隆中的视觉域鲁棒性

Enhancing Visual Domain Robustness in Behaviour Cloning via Saliency-Guided Augmentation

Zheyu Zhuang, Ruiyu Wang, Nils Ingelhag, Ville Kyrki, Danica Kragic

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对行为克隆中常规增强方法应对视觉域偏移效果差的问题,提出RoboSaGA显著性引导增强方法,提升了模型的视觉域鲁棒性且保留了域内性能。

中文摘要 AI 辅助

在基于视觉的行为克隆(BC)中,随机裁剪、颜色抖动等常规图像增强方法在存在阴影、干扰项、背景变化等大幅视觉域偏移时往往效果不佳。叠加式增强方法通过混合域内和域外图像,在计算机视觉中展现出提升泛化能力的潜力,但其是否适用于BC仍不确定,因为必须保留任务关键语义、时空关系以及智能体-目标交互。为解决该问题,我们提出RoboSaGA,一种叠加式家族内的显著性引导增强方法,专为基于视觉的BC定制。RoboSaGA利用策略驱动的显著性在像素级别动态调整增强强度,在任务无关区域进行激进增强,同时保留任务关键信息。它可无缝集成到现有架构中,无需结构修改或额外学习目标。在仿真和真实场景中的实验表明,RoboSaGA在保留域内性能的同时,大幅提升了对视觉域偏移的鲁棒性,包括干扰项、背景变化以及光照和阴影变化。代码可在该https URL获取。

英文摘要

In vision-based behavior cloning (BC), conventional image augmentations such as Random Crop and Color Jitter often fall short under substantial visual domain shifts, including changes in shadows, distractors, and backgrounds. Superimposition-based augmentations, which blend in-domain and out-of-domain images, have shown promise for improving generalization in computer vision, but their suitability for BC remains uncertain because task-critical semantics, spatiotemporal relationships, and agent-target interactions must be preserved. To address this, we introduce RoboSaGA, a Saliency-Guided Augmentation method within the superimposition family tailored for vision-based BC. RoboSaGA dynamically adjusts augmentation intensity at the pixel level using policy-driven saliency, enabling aggressive augmentation in task-irrelevant regions while preserving task-critical information. It integrates seamlessly into existing architectures without requiring structural modifications or additional learning objectives. Experiments in both simulated and real-world settings show that RoboSaGA preserves in-domain performance while substantially improving robustness to visual domain shifts, including distractor and background changes, as well as lighting and shadow variations. Code is available at https://github.com/Zheyu-Zhuang/RoboSaGA.

发表机构

  • KTH Royal Institute of Technology(皇家理工学院)
  • Aalto University(阿尔托大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑