arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

免训练行为克隆

Training-free Behavior Cloning

Maximilian Adang, Timothy Chen, Lars Osterberg, Aiden Swann, Mac Schwager

arXiv 2609.30134首次发表:更新:

发表机构

Stanford University(斯坦福大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出行为预测控制(BPC),一种无需端到端训练的行为克隆方法,通过动作感知检索、Hankel先验和残差校正,在保持竞争力的同时实现快速策略合成与可追溯预测。

AI 中文摘要

神经行为克隆将演示压缩到大型模型中,使得单个动作难以追踪且策略更新成本高昂。检索策略保留了对演示的访问,但难以应对记录行为与实时行为之间的不匹配。我们引入了行为预测控制(BPC),它通过结合动作感知检索度量、基于Hankel的动作延续先验和闭式一步残差校正,在不进行端到端策略训练的情况下合成策略。受行为系统理论启发,BPC通过融合最能重建近期运行时观测-动作历史存储的观测-动作数据来预测未来动作。在模拟基准和真实机器人部署中,BPC与学习策略(如π₀.₅)相比具有竞争力(在某些情况下超越它),同时将策略拟合时间从数小时缩短到消费级GPU上的数秒,并支持在Jetson Orin Nano上以超过75Hz的频率进行闭环控制。检索到的演示窗口及其系数还提供了任务进展的内在估计。在部署的策略中保留演示,使其预测可追溯到支持轨迹,并通过演示库实现行为修正。

英文摘要

Neural behavior cloning compresses demonstrations into large models, making individual actions difficult to trace and policy updates costly. Retrieval policies retain access to demonstrations but struggle with mismatch between recorded and live behavior. We introduce Behavior Predictive Control (BPC), which synthesizes policies without end-to-end policy training by combining an action-aware retrieval metric, a Hankel-based action-continuation prior, and a closed-form one-step residual correction. Inspired by behavioral systems theory, BPC predicts future actions by blending stored observation-action data that best reconstructs the recent runtime observation--action history. Across simulated benchmarks and real-robot deployments, BPC is competitive with learned policies such as $π_{0.5}$ (surpassing it in some cases), while reducing policy fitting from hours to seconds on consumer GPUs and supporting closed-loop control upwards of 75 Hz on a Jetson Orin Nano. The retrieved demonstration windows and their coefficients also provide an intrinsic estimate of task progress. Retaining demonstrations within the deployed policy makes its predictions traceable to supporting trajectories and enables behavior revision through the demonstration bank.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑