TimelyDAgger:面向VLA策略改进的时序感知专家查询
TimelyDAgger: Timing-Aware Expert Querying for VLA Policy Improvement
AI总结:
TimelyDAgger通过结合Bridge-PCA特征监控与反馈引导阈值自适应,优化机器人门控DAgger的专家接管时机,提升VLA策略训练效果,并在多数设置下实现更高成功率。
AI中文摘要:
DAgger通过聚合策略执行过程中访问状态下的专家监督来改进机器人策略。机器人门控DAgger自动化专家查询,允许机器人决定何时请求专家接管。虽然现有门控强调检测援助需求,但接管时机也塑造了这些演示的内容及其对策略学习的价值。我们提出TimelyDAgger,结合Bridge-PCA对内部视觉-语言-动作(VLA)特征的监控,以及基于专家行为的反馈引导阈值自适应,以改进接管时机。我们引入了一个评估框架,将失败检测、接管时机和策略改进联系起来,包括目标对齐监督比率(TASR),用于在不重新训练的情况下评估监督质量。实验表明,接管时机影响策略学习,在匹配专家动作预算下,TimelyDAgger在大多数评估设置中实现了具有竞争力的失败检测和更高的训练后成功率。
英文摘要:
DAgger improves robot policies by aggregating expert supervision from states visited during policy execution. Robot-gated DAgger automates expert queries, allowing the robot to decide when to request expert takeover. While existing gates emphasize detecting the need for assistance, takeover timing also shapes the content of these demonstrations and their value for policy learning. We propose TimelyDAgger, combining Bridge-PCA monitoring of internal vision-language-action (VLA) features with Feedback-guided Threshold Adaptation based on expert behavior to improve takeover timing. We introduce an evaluation framework linking failure detection, takeover timing, and policy improvement, including Target-Aligned Supervision Ratio (TASR) for assessing supervision quality without retraining. Experiments show that takeover timing affects policy learning, with TimelyDAgger achieving competitive failure detection and higher post-training success in most evaluated settings under matched expert-action budgets. Project website: https://seen-e.github.io/TimelyDagger/.