arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.23831cs.ROcs.LG

等待时学习行动:考虑推理延迟的通用机器人策略的强化学习微调

Learning to Act While Waiting: RL Finetuning of Generalist Robot Policies Under Inference Latency

Brian Zhu, Momen Khalil, E Harrison, Emanuele Poggi, Philipp Schmitt, Bernd Kast, Philine Meister, Pranav Atreya, Qiyang Li, Finn Ferchau, Cesar Colmenero, Yash… 展开作者

Brian Zhu, Momen Khalil, E Harrison, Emanuele Poggi, Philipp Schmitt, Bernd Kast, Philine Meister, Pranav Atreya, Qiyang Li, Finn Ferchau, Cesar Colmenero, Yash Shahapurkar, Gokul Narayanan, Melih Erdogan, Kai Wurm, Georg von Wichert, Oier Mees, Eugen Solowjow, Andrew Wagenmaker, Sergey Levine

首次发表
浏览论文内容

中文总结 AI 辅助

本研究针对通用机器人策略因推理延迟破坏马尔可夫假设导致标准RL失效的问题,提出ARLI框架,通过状态增强等设计实现延迟下的RL微调,在模拟及真实任务中表现优异。

中文摘要 AI 辅助

虽然强化学习(RL)可让通用机器人策略在部署期间持续改进,但现代通用策略(如视觉语言动作模型VLAs)的庞大模型规模,对RL有效改进构成了根本障碍。尤其是其严重的推理延迟——可能导致机器人动作停顿或不流畅——会改变有效环境动态,若未正确考虑,将破坏RL所依赖的马尔可夫假设,导致标准RL算法完全失效。本研究提出一种感知延迟的框架:带中间信息的异步强化学习(ARLI),可在推理延迟下实现通用策略的RL改进。该框架基于异步推理方法,通过将动作生成与执行交错以隐藏延迟,同时解决其与RL的不兼容性,具体通过两项贡献设计低延迟RL策略,在推理窗口内最大化反应性:一是状态增强,通过纳入已执行动作和推理过程中的观测值,恢复近似马尔可夫结构;二是相关技术优化。我们在模拟和真实世界的操作任务中评估该方法,发现其能在标准RL完全失效的推理延迟场景下实现有效微调,甚至在理想无延迟设置中达到或超过标准RL的性能。

英文摘要

While reinforcement learning (RL) allows generalist robot policies to continually improve during deployment, the large model size of modern generalist policies, such as VLAs, poses a fundamental obstacle to effective RL improvement. In particular, their severe inference latency---which can lead to pauses or jerky movements---can alter the effective environment dynamics and, if not correctly accounted for, break the Markov assumption that RL relies on, causing standard RL algorithms to fail completely. In this work, we introduce a latency-aware framework, Asynchronous RL with Intermediate Information (ARLI), that enables RL-based improvement of generalist policies under inference delays. Our framework builds on asynchronous inference approaches, which interleave action generation with execution to hide latency, and addresses its incompatibility with RL by providing a low-latency RL policy design that maximizes reactivity within the inference window through two contributions: state augmentations that restore near-Markovian structure by incorporating committed actions and a mid-inference observation. We evaluate our approach across simulated and real-world manipulation tasks, and find that it enables effective finetuning under inference delays where standard RL fails entirely, even matching or exceeding the performance of standard RL in idealized no-latency settings.

发表机构

  • Siemens(西门子)
  • UC Berkeley(加州大学伯克利分校)
  • Microsoft(微软)
  • ETH Zurich(苏黎世联邦理工学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑