发表机构
Northwestern; CMU; UIUC; UCSD(西北大学; 卡内基梅隆大学; 伊利诺伊大学厄巴纳-香槟分校; 加州大学圣迭戈分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出Video2Skill基准,研究从流式视频中自动发现可复用技能,发现现有VLM在技能分组与库扩展上存在瓶颈,并引入CLaRe方法部分改善。
AI 中文摘要
操作行为在物体和场景之间差异很大,但它们共享一小部分可复用的技能,利用这些技能进行规划有助于具身智能体泛化到新任务。然而,智能体只能使用它已知的技能进行规划。从观察到的经验中恢复技能(即规划的逆过程)会随时间积累这一知识,并为训练未来的智能体提供技能数据。视觉语言模型(VLMs)能很好地描述单个操作事件,但它们能否将一系列事件组织成可复用的技能?我们将此问题定义为流式具身技能发现(SESD):模型按顺序观看视频,并维护一个持续更新的技能库,该库会影响其后续决策。为了系统性地衡量这一能力,我们引入了Video2Skill基准,涵盖机器人桌面操作和人类厨房活动,并测试三项核心能力:(i)在时间上定位操作事件,(ii)将相同变换的事件分组,以及(iii)决定何时复用现有技能或创建新技能。在19个开源VLM中,许多模型对事件的分组接近随机水平,且模型规模并未持续带来改进。它们的错误取决于感知与技能库更新的耦合方式:联合模型将不同的变换合并为一个技能,而从文本描述更新技能库的模型则会重复出现相似技能。监督微调(包括我们提出的反事实技能库状态再平衡方法CLaRe)改善了分组,但暴露了更深的瓶颈:训练后的模型能巩固熟悉的技能,却很少扩展技能库。它们的技能库规模停滞在参考规模的一半以下,训练中未见过的变换虽能在时间上被定位,但几乎从未被赋予新技能。因此,识别现有技能何时不足成为核心挑战。
英文摘要
Manipulation behaviors vary widely across objects and scenes, but they share a small set of reusable skills, and planning with these skills helps embodied agents generalize to new tasks. Yet an agent can only plan with skills it knows. Recovering skills from observed experience, the inverse of planning, builds this knowledge over time and yields skill data for training future agents. Vision-Language Models (VLMs) describe individual manipulation events well, but can they organize a stream of events into reusable skills? We formulate this problem as Streaming Embodied Skill Discovery (SESD): a model watches videos in sequence and maintains a persistent skill library that shapes its later decisions. To systematically measure this ability, we introduce Video2Skill, a benchmark that covers robot tabletop manipulation and human kitchen activity and tests three core capabilities: (i) locating manipulation events in time, (ii) grouping events of the same transformation, and (iii) deciding when to reuse an existing skill or create a new one. Across 19 open-source VLMs, many models group events at near-chance level, and scale does not consistently help. Their errors depend on how perception and library updates are coupled: joint models merge distinct transformations into one skill, while models that update the library from text descriptions duplicate recurring ones. Supervised fine-tuning, including our counterfactual library-state rebalancing (CLaRe), improves grouping but exposes a deeper bottleneck: trained models consolidate familiar skills yet rarely expand the library. Their libraries stall below half the reference size, and transformations unseen in training are located in time but almost never given a new skill. Recognizing when existing skills are insufficient thus emerges as the central challenge.
CommentsProject page: https://andyzworks.github.io/video2skill/