发表机构
Karlsruhe Institute of Technology; University of Leeds(卡尔斯鲁厄理工学院; 利兹大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对视觉模仿学习中RGB观测缺乏几何线索导致策略对相机扰动脆弱的问题,提出RayViT架构,结合辅助余弦相似度损失,在RoboCasa基准及真实机器人任务中提升了策略的鲁棒性与多任务表现。
AI 中文摘要
视觉模仿学习使机器人能够直接从图像中获取视觉运动技能,但RGB观测缺乏明确的几何线索,导致学习到的策略对相机扰动较为脆弱。为解决该问题,我们提出光线条件视觉Transformer编码器(Ray-conditioned Vision Transformer Encoder,RayViT),这是一种将相机几何信息注入预训练ViT骨干网络的轻量型架构。RayViT将相机几何表示为普吕克光线图(Plücker ray map),将其分块为光线特征,并使用门控交叉注意力生成光线条件类token。这些光线特征作为密集位置嵌入添加,而光线类token则替换原始ViT类token以提供几何感知的总结表示。我们将该方法与辅助余弦相似度损失相结合,以持续提升几何感知token的性能与鲁棒性。在仿真和真实机器人任务上的实验表明,RayViT在多任务RoboCasa基准测试中,相机扰动下的鲁棒性较基线提升约13个百分点,在真实世界多任务成功率中,平均完成阶段数较基线提升1.78。
英文摘要
Visual imitation learning enables robots to acquire visuomotor skills directly from images, yet RGB observations lack explicit geometric cues, making learned policies brittle to camera perturbations. To address this, we propose \textbf{Ray-conditioned Vision Transformer Encoder (RayViT)}, a lightweight architecture that injects camera geometry into pretrained ViT backbones. RayViT represents camera geometry as a Plücker ray map, patchifies it into ray features, and uses gated cross-attention to produce a ray-conditioned class token. These ray features are added as dense positional embeddings, while the ray class token replaces the original ViT class token to provide a geometry-aware summary representation. We combine this approach with an auxiliary cosine similarity loss to consistently improve the performance and robustness for geometry-aware tokens. Experiments on sim- and real-robot tasks demonstrate that RayViT improves robustness by approximately 13 percentage points under camera perturbations in multi-task RoboCasa benchmark and by 1.78 average completed stages in real-world multi-task success rate compared to baselines.