arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于智能体机器人的递归视频上下文学习

Recursive Video In-Context Learning for Agentic Robot

Wenrui Bao, Xinxin Liu, Bingxin Xu, Yuzhang Shang

arXiv 2610.06843首次发表:更新:

发表机构

University of Central Florida; University of Southern California(中佛罗里达大学; 南加州大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出RV-ICL,一种无需训练的递归视频上下文学习方法,将演示视频组织为层级结构供智能体按需访问,在LIBERO基准上将成功率提升至96.5%和95.8%。

AI 中文摘要

编排冻结的视觉-语言-动作(VLA)策略的LLM智能体通过文本记忆在多个回合中提升表现,文本记忆记录了智能体做了什么,但未记录任务如何完成。演示视频展示了任务如何完成,但难以适配智能体的上下文。完整视频拖慢每一轮交互,固定关键帧丢失了决定抓取是否成功的接触细节,而智能体的需求从规划时的任务结构转变为每个接触点周围的帧。我们提出递归视频上下文学习(RV-ICL),一种无需训练的方法,将演示转化为智能体导航的层级结构,而非接收的提示。该层级结构由演示的子事件(如抓取和释放)构建。其层级逐渐细化,从整个任务的关键帧到阶段、时刻和短视频片段,并通过只读工具暴露。智能体在规划前读取粗粒度层级。执行过程中,每当步骤需要更多细节时,它重新进入层级,仅加载当前子目标的片段。每个任务只需一个演示即可。基于RPent,RV-ICL在LIBERO-PRO上将成功率从92.6%提升至96.5%,在LIBERO-Plus上从86.7%提升至95.8%。

英文摘要

LLM agents that orchestrate frozen vision-language-action (VLA) policies improve across episodes through text memory, which records what the agent did but not how the task is done. A demonstration video shows it, but fits poorly into an agent's context. The full video slows every turn, fixed keyframes lose the contact detail that decides whether a grasp holds, and what the agent needs shifts from the task's structure while planning to the frames around each contact. We introduce Recursive Video In-Context Learning (RV-ICL), a training-free method that turns a demonstration into a hierarchy the agent navigates rather than a prompt it receives. The hierarchy is built from the sub-events of the demonstration, such as grasps and releases. Its levels grow finer, from keyframes of the whole task to phases, moments and short clips, and are exposed through read-only tools. The agent reads the coarse levels before planning. During execution it re-enters the hierarchy whenever a step needs more detail and loads only the clip of its current sub-goal. One demonstration per task is enough. Built on RPent, RV-ICL raises success from 92.6% to 96.5% on LIBERO-PRO and from 86.7% to 95.8% on LIBERO-Plus.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑