arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

HALO-WA:用于世界-动作模型的混合注意力潜在引导在线强化学习

HALO-WA: Hybrid-Attention Latent-Guided Online Reinforcement Learning for World-Action Models

Angen Ye, Weijie Ke, Xiaofeng Wang, Xinze Chen, Chaojun Ni, Guosheng Zhao, Boyuan Wang, Zheng Zhu, Junjie Xie, Dapeng Zhang

arXiv 2607.04265首次发表:更新:

发表机构

Institute of Automation, Chinese Academy of Sciences; School of Artificial Intelligence, University of Chinese Academy of Sciences; GigaAI; Tsinghua University(中国科学院自动化研究所; 中国科学院大学人工智能学院; 极佳科技; 清华大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对世界-动作模型在现实精度任务中易受多种误差影响的问题,提出HALO-WA框架,利用潜在特征和动作先验,通过混合注意力结构实现快速在线适应,提升动作块精度,实验验证效果显著。

AI 中文摘要

世界-动作(WA)模型可为通用机器人操作生成长时动作块,但在现实精度任务中易受校准、感知和接触动力学误差影响,常在对齐或插入的最后几毫米失败。我们提出HALO-WA,一种用于WA模型的混合注意力潜在引导在线强化学习框架,通过轻量级actor-critic适配器利用WA生成过程中的潜在特征和动作先验,实现对实际部署误差的快速在线适应。HALO-WA引入混合注意力结构,在根据视觉上下文和末期校正要求从WA潜变量中读取任务相关信息时保持动作块的时间一致性,从而产生精细的动作块。我们在四个现实世界精度操作任务上验证了HALO-WA,将WA-base的平均成功率从26.4%提高到87.1%,比最强基线高出19.2个百分点,同时每个任务仅需45 - 75分钟的在线训练。为便于重现,我们在RoboTwin中进行了补充模拟实验,并在这个https URL上发布了代码。

英文摘要

World-action (WA) models can generate long-horizon action chunks for general-purpose robotic manipulation, but they remain vulnerable to calibration, perception, and contact-dynamics errors in real-world precision tasks, often failing in the final few millimeters of alignment or insertion. We propose HALO-WA, a hybrid-attention latent-guided online reinforcement learning (RL) framework for WA models, which leverages latent features and action priors from the WA generation process through a lightweight actor-critic adapter to enable fast online adaptation to real deployment errors. HALO-WA introduces a hybrid-attention structure that preserves the temporal consistency of action chunks while reading task-relevant information from WA latents conditioned on visual context and end-stage correction requirements, thereby producing refined action chunks. We validate HALO-WA on four real-world precision manipulation tasks, where it improves the average success rate from 26.4\% for WA-base to 87.1\%, outperforming the strongest baseline by 19.2 percentage points while requiring only 45--75 minutes of online training per task. To facilitate reproducibility, we further conduct supplementary simulation experiments in RoboTwin and release the code at https://github.com/YeanRoot/HALO-WA.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑