arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向自主挖掘的视觉引导目标条件控制

Vision Guided Target Conditioned Control for Autonomous Excavation

Shuai Zhao, Ji-An Pan, Junwei Li, Xun Tang, Fansen Xi, Qing Xu, Keqiang Li, Jianqiang Wang

arXiv 2608.21778首次发表:更新:

发表机构

Liaoning University; Northeastern University; Tsinghua University(辽宁大学; 东北大学; 清华大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对自主挖掘问题,提出结合视觉目标条件与配对演示的掩码条件动作块Transformer框架,在模拟任务中显著提升挖掘成功率与堆料清理效率,为挖掘机自动化提供实用方案。

AI 中文摘要

自主挖掘需要一套智能控制系统,该系统能够在接触密集型土壤交互下,将空间作业意图转化为协调的铲斗运动。本文提出一种用于自主挖掘的目标条件智能控制框架,该框架基于物理可变形土壤模拟工作流构建。图像对齐的目标掩码作为期望挖掘区域的视觉空间指令,而掩码条件动作块Transformer(mask-conditioned Action Chunking Transformer)则将多视角RGB观测、本体感知数据以及目标掩码映射为时间扩展的操纵杆指令。为减少忽略目标的行为,演示采用配对条件监督组织,即对相同或高度匹配的场景,使用不同的目标掩码及对应的动作块进行演示。该框架通过诊断操作任务以及包含单铲和连续堆料清理协议的挖掘模拟基准进行评估:在操作任务中,无条件ACT的目标成功率为4%,非配对掩码条件ACT的成功率为63%,配对掩码条件ACT的成功率为96%;在连续堆料清理任务中,配对掩码条件ACT可移除76.8%的堆料,而两个基准方法的移除率分别为27.4%和15.7%,同时其人类归一化效率达91.0%。结果表明,视觉目标条件、配对演示结构以及动作块控制构成了一套实用的挖掘机自动化网络物理模拟管线。

英文摘要

Autonomous excavation requires an intelligent control system that can convert spatial work intent into coordinated bucket motion under contact-rich soil interaction. This paper presents a target-conditioned intelligent control framework for autonomous excavation in a physics-based deformable-soil simulation workflow. An image-aligned target mask serves as a visual spatial command for the desired digging region, while a mask-conditioned Action Chunking Transformer maps multi-view RGB observations, proprioception, and the target mask to temporally extended joystick commands. To reduce target-ignoring behavior, demonstrations are organized with paired-condition supervision, where the same or closely matched scene is demonstrated with different target masks and corresponding action chunks. The framework is evaluated through both a diagnostic manipulation task and an excavation simulation benchmark with single-scoop and sequential pile-clearing protocols. In manipulation, target success is 4\% for no-condition ACT, 63\% for non-paired mask-conditioned ACT, and 96\% for paired-condition mask-conditioned ACT. In sequential pile clearing, paired-condition mask-conditioned ACT removes 76.8\% of the pile versus 27.4\% and 15.7\% for the two baselines, with 91.0\% human-normalized efficiency. The results show that visual target conditioning, paired demonstration structure, and action-chunk control form a practical cyber-physical simulation pipeline for excavator automation.

Comments5 pages, 3 figures, 3 tables. Accepted at ISCSIC 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑