arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

观察、回忆、行动:并发具身流中的常开机器人

Watch, Recall, Act: Always-On Robots in Concurrent Embodied Streams

Ding Yi, Peiwen Sun, Chenchu Rong, Jianan Wang, Xili Dai, Xiangyu Yue, Xi Lin

arXiv 2609.28429首次发表:更新:

发表机构

Shanghai Jiaotong University; The Chinese University of Hong Kong; Nanjing University; Astribot; Juxi Tech(上海交通大学; 香港中文大学; 南京大学; Astribot; 聚溪科技)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对常开机器人在并发具身流中的挑战,提出ARMS策略,利用预训练骨干加三个轻量模块异步处理实时感知、具身状态和自历史,在ARMS数据集训练后达到45%成功率,显著优于基线。

AI 中文摘要

常开机器人面临的是永不重置的无尽流:指令到来又失效,场景不断变化,其自身过往行动也重塑着它必须推理的内容。当今的动作模型是为相反情况而构建的:固定指令、无任务中干预、单步推理。在开放世界环境中,机器人必须观察实时流以获取远期线索,回忆自身久远的过往行动,并在双臂并发条件下执行这些行动。我们提出ARMS(多模态流中的常开机器人),一种刻意简化的流式策略:一个预训练的π0.5骨干网络,辅以三个轻量级模块,将实时感知、具身状态和机器人自身过往行动转化为骨干网络在执行前读取的上下文。这些模块异步更新上下文,因此观察和回忆从不阻塞执行,且双臂同时动作。ARMS并非发明新机制,而是将这些学习到的上下文提供者与记录哪只手臂在何时做了什么的智能体因果自历史相集成。为无需额外标注地监督它们,我们构建了ARMS数据集,其分阶段构建脚本本身即从真实双臂遥操作中标注每个模块。在该数据集上训练后,ARMS在组合任务上达到45%的成功率,而我们四个主要基线中最强者仅为28%,消融实验证实记忆模块、具身状态头和异步并发均不可或缺。

英文摘要

An always-on robot faces an endless stream that never resets: instructions arrive and lapse, the scene changes, and its own past actions reshape what it must reason about. Today's action models are built for the opposite: a fixed instruction, no mid-task intervention, single-step reasoning. In an open-ended world a robot must watch a live stream for far-future cues, recall its own far-past actions, and act on them under dual-arm concurrency. We present ARMS (Always-on Robot in Multi-modal Streams), a deliberately simple streaming policy: a single pretrained $π$0.5 backbone augmented by three lightweight modules that turn live perception, embodied states, and the robot's own past actions into context the backbone reads before it acts. The modules update this context asynchronously, so watching and recalling never block acting and the two arms act at once. Rather than inventing new mechanisms, ARMS integrates these learned context providers with an agent-causal self-history that logs which arm did what, and when. To supervise them without extra annotation, we build ARMS Dataset, whose staged construction script itself labels every module from real dual-arm teleoperation. Trained on it, ARMS reaches 45% on the combined task against 28% for the strongest of our four main baselines, and ablations confirm the memory module, the embodied-state head, and asynchronous concurrency are each necessary.

Comments11 pages, 3 figures. Accepted to the 10th Conference on Robot Learning (CoRL 2026)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑