arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.27780cs.ROcs.CV

任务原型引导的流匹配用于视觉-语言机器人操作中的少样本泛化

Task-Prototype Guided Flow Matching for Few-Shot Generalization in Vision-Language Robot Manipulation

  • Henan Institute of Science and Technology(河南科技学院)

机构由 AI 辅助整理,请以论文原文为准。

Yizhao Wang, Guantao Zhang, Jingbo Wang

AI总结:

本文提出TP-Flow框架,利用任务原型引导流匹配,通过少量演示实现视觉-语言机器人操作的少样本泛化,在多个基准上显著提升成功率并保持实时性能。

AI中文摘要:

视觉-语言机器人操作策略能够遵循语义指令,但仅凭少量演示将其适应到新流程仍然困难,因为语言无法充分指定接触时机、运动阶段、纠正行为和执行风格。本文提出了任务原型引导的流匹配(TP-Flow),一种少样本操作框架,它将支持演示转换为结构化的任务原型标记,并用它们来引导初始流先验和速度场。TP-Flow采用带有可学习查询的对称交叉注意力来提取阶段级原型,参数化任务自适应的初始分布,并通过门控自适应归一化注入原型信息。它使用情节式支持-查询目标和原型对比正则化进行训练,因此在训练期间模拟了少样本适应,同时抑制了干扰信息。在LEROBOT-ARM-SO101平台上,TP-Flow在1次、4次和6次样本设置下分别实现了66.8%、79.6%和82.1%的成功率,少样本AUC为75.5%。在1次样本下,它比CFM、Pooled-Demo CFM和In-Context Flow分别提高了29.8、14.5和9.9个百分点。它还在新物体迁移、目标重组、长时程组合和接触/纠正任务中提高了对保留目标组的泛化能力。TP-Flow通过六个在线原型标记、54.3毫秒延迟、3.9 GB峰值内存和10 Hz控制率保持实时执行,同时将噪声支持下的成功率下降降至6.2%。理论诊断表明,原型距离与动作分布距离一致,自适应先验降低了传输成本,门控调制使测得的轨迹偏差保持在导出的ODE界限以下。代码仓库因匿名评审而省略。

英文摘要:

Vision-language robot manipulation policies can follow semantic instructions, but adapting them to a new procedure from only a few demonstrations remains difficult because language underspecifies contact timing, motion phases, corrective behavior, and execution style. This paper presents Task-Prototype Guided Flow Matching (TP-Flow), a few-shot manipulation framework that converts support demonstrations into structured task-prototype tokens and uses them to guide both the initial flow prior and the velocity field. TP-Flow employs symmetric cross-attention with learnable queries to extract phase-level prototypes, parameterizes a task-adaptive initial distribution, and injects prototype information through gated adaptive normalization. It is trained with an episodic support-query objective and prototype contrastive regularization, so few-shot adaptation is simulated during training while nuisance information is suppressed. On the LEROBOT-ARM-SO101 platform, TP-Flow achieves 66.8\%, 79.6\%, and 82.1\% success rates under 1-, 4-, and 6-shot settings, with a 75.5\% few-shot AUC. At 1-shot, it improves over CFM, Pooled-Demo CFM, and In-Context Flow by 29.8, 14.5, and 9.9 percentage points. It also improves held-out target-group generalization across novel-object transfer, goal recombination, long-horizon composition, and contact/correction tasks. TP-Flow maintains real-time execution with six online prototype tokens, 54.3 ms latency, 3.9 GB peak memory, and a 10 Hz control rate, while reducing the noisy-support success drop to 6.2\%. Theoretical diagnostics show that prototype distance aligns with action-distribution distance, the adaptive prior reduces transport cost, and gated modulation keeps measured trajectory deviations below the derived ODE bound. The code repository is omitted for anonymous review.

补充信息

↑