发表机构
Samsung Robotics eXperience; Shanghai Jiao Tong University; Samsung Research(三星机器人体验中心; 上海交通大学; 三星研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
RoboICL通过分离演示上下文与交互记忆,利用上下文学习提升GPT-6 Astra在机器人控制中的精度和长时程性能,在RoboDojo基准上显著超越零样本基线。
AI 中文摘要
通用视觉语言模型为零样本机器人控制提供了一种有前景的途径:GPT-6 Astra在开放式以及语言或图像条件化的操作任务中表现出色,但在高精度和长时程任务上仍明显较弱。我们提出了RoboICL,一个无需机器人特定参数更新或学习型VLA的上下文机器人控制框架,以缩小这些差距。RoboICL将“演示上下文”(在可用时提供记录示例)与“交互记忆”(累积模型自身动作和观察结果)分离。两者均使用共享的“观察-动作-接收-观察”语法。为跨任务阶段保留经验,RoboICL结合了采样演示块和有界锚定记忆。固定锚点使早期回滚交互可用于上下文学习,而最新交互支持即时错误纠正。在30个RoboDojo任务中,除Open任务使用零样本外,其余任务均使用一个演示,RoboICL在每个类别上比官方零样本GPT-6 Astra提高了20至27个进度分数点。它在Memory和Open任务上领先排行榜基线,在Precision任务上达到与最强基线相当的性能,并在Long-Horizon任务上保持竞争力。其30任务总体得分为50.64,而最强基线为33.68。在单独的10任务子集上,RoboICL得分为60.60,与π0.5 + GPT-6 Astra混合方法相差2.00分以内。在三个真实机器人任务上,平均进度从零样本的14.45提高到单样本的63.33和三次样本的78.89。在两个开发任务上,可选的Jev门控动作复用将GPT-6 Astra调用减少了33%至48%。代码可在以下网址获取:this https URL。
英文摘要
General-purpose vision-language models offer a promising way to zero-shot robot control: \gptastra{} excels at open-ended and language- or image-conditioned manipulation but remains substantially weaker on high-precision and long-horizon tasks. We introduce \emph{RoboICL}, an in-context robot-control framework that narrows these gaps without robot-specific parameter updates or a learned VLA. RoboICL separates \emph{demonstration context}, which provides recorded examples when available, from \emph{interaction memory}, which accumulates the model's own actions and observed outcomes. Both use a shared observation--action--receipt--observation grammar. To preserve experience across task stages, RoboICL combines sampled demonstration blocks with bounded anchored memory. Fixed anchors keep earlier rollout interactions available for in-context learning, while the latest interaction supports immediate error correction. Across 30 RoboDojo tasks, using zero shot for Open and one demonstration elsewhere, RoboICL improves on official zero-shot \gptastra{} by 20--27 progress-score points in every category. It leads the leaderboard baselines on Memory and Open, achieves comparable performance to the strongest Precision baseline, and remains competitive on Long-Horizon. Its 30-task Overall score is 50.64, versus 33.68 for the strongest baseline. On a separate ten-task subset, RoboICL scores 60.60, within 2.00 points of the $π_{0.5}$ + \gptastra{} hybrid approach. On three real-robot tasks, mean progress rises from 14.45 at zero shot to 63.33 at one shot and 78.89 at three shots. On two development tasks, optional Jev-gated action reuse reduces \gptastra{} calls by 33--48\%. Code is available at \href{https://github.com/Mosi-AI/RoboICL}{https://github.com/Mosi-AI/RoboICL}.