arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

双过程原子技能学习:解耦语义推理与实时控制

Dual-Process Atomic Skill Learning: Decoupling Semantic Reasoning and Real-Time Control

Jun Chen, Erdemt Bao, Wenlong Dong, Jierui Liu, Qi Cai, Hao Wan, Shaopeng Li, Weijun Qin, Jing Liang, Huiping Zhuang

arXiv 2607.10625首次发表:更新:

发表机构

University of Electronic Science and Technology of China; Huazhong University of Science and Technology; Southern University of Science and Technology; South China University of Technology; EbTech Co. Ltd.(电子科技大学; 华中科技大学; 南方科技大学; 华南理工大学; 依比特科技有限公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对语言条件模仿学习推广到多步组合任务的挑战,提出双过程原子技能学习框架DASL,通过解耦语义推理与实时控制,含低频和高频策略,有效减轻技能码本干扰问题,性能显著优于现有基线。

AI 中文摘要

语言条件模仿学习对使机器人按自然语言指令执行复杂任务至关重要,但推广到多步组合任务仍是重大挑战。现有分层方法因联合训练中高层技能推理与低层动作生成紧密耦合,常面临训练不稳定和码本崩溃问题。受认知双过程理论启发,提出双过程原子技能学习(DASL),一种新颖的异步分层模仿学习框架,将缓慢的语义推理与快速实时运动控制解耦。DASL包括通过矢量量化预测可解释离散技能的低频策略,以及利用潜在扩散模型和决策变换器根据这些潜在技能生成精确动作的高频策略。通过异步协调这些模块并利用扩散构建潜在空间,减轻了联合训练范式中常见的技能码本干扰问题。跨模拟基准和实验评估表明,DASL显著优于现有基线,在技能获取和对未见指令的组合泛化方面表现出色。

英文摘要

Language-conditioned Imitation Learning (IL) is essential for enabling robots to perform complex tasks following natural language instructions. However, generalizing to multi-step compositional tasks remains a significant challenge. While hierarchical approaches attempt to address this by decomposing tasks into atomic skills, existing methods often suffer from training instability and codebook collapse due to the tight coupling between high-level skill reasoning and low-level action generation in joint training paradigms. Inspired by the Dual-Process Theory of cognition, we propose Dual-Process Atomic Skill Learning (DASL), a novel asynchronous hierarchical imitation learning framework that decouples slow semantic reasoning from fast, real-time motion control. DASL comprises a Slow-Frequency Policy that predicts interpretable, discrete skills via Vector Quantization, and a High-Frequency Policy that leverages a latent diffusion model and a Decision Transformer to generate precise actions conditioned on these latent skills. By asynchronously coordinating these modules and utilizing diffusion to structure the latent space, our framework mitigates the skill codebook interference problem common in joint training paradigms. Evaluations across simulation benchmarks and experiment demonstrate that DASL significantly outperforms state-of-the-art baselines, excelling in skill acquisition and compositional generalization to unseen instructions. GitHub page: https://github.com/Hatakekaka/DASL

Comments28 pages,20 figures,21 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑