arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.10021cs.RO

RoboDrop:通过局部梯度兼容性策展VLA后训练数据

RoboDrop: Curating VLA Post-Training Data via Local Gradient Compatibility

Runze Xu, Yuanfan Xu, Cuijie Xu, Shuang Dai, Yining Li, Yu Wang, Jincheng Yu

首次发表
浏览论文内容

中文总结 AI 辅助

RoboDrop通过沿训练轨迹的局部梯度兼容性审计监督,自动筛选机器人后训练数据,提升策略性能,真实机器人成功率从35.0%升至67.5%。

中文摘要 AI 辅助

视觉-语言-动作(VLA)模型通过大规模预训练获得广泛泛化能力,但将其适应于新任务和机器人本体仍需要对新收集的数据进行后训练。与预训练不同,后训练针对任务和本体特定的适应,因此对数据质量尤为敏感。在实践中,收集的机器人数据集通常包含异构错误,包括执行错误、传感器漂移和时间戳错位,这些错误可能损害后训练和策略性能。人工检查成本高昂,而现有的数据清洗方法通常针对特定损坏类型定制。为应对这些挑战,我们引入了RoboDrop,一个数据策展框架,它使用沿训练轨迹测量的局部梯度兼容性作为其对后训练性能影响的代理来审计监督。在一次epoch的热身运行中,RoboDrop通过将每个候选样本的梯度与任务语义和视觉匹配的验证样本的梯度进行比较,在线对每个候选样本进行评分。得到的样本分数在情节级别聚合,一个简单的自动后处理规则将其转换为过滤决策。我们在受控的观测-动作损坏、模拟中的自然次优演示以及包含非专家收集错误的真实机器人数据集上评估了RoboDrop。在这些设置中,RoboDrop比先前方法更准确地区分不可靠演示,而在策展数据上进行后训练持续产生更强的下游策略,平均真实机器人部署成功率从35.0%提升到67.5%。这些结果确立了训练轨迹感知、上下文条件监督审计作为稳健VLA后训练的有效方法。

英文摘要

Vision--language--action (VLA) models acquire broad generalization through large-scale pretraining, yet adapting them to a new task and robot embodiment still requires post-training on newly collected data. Unlike pretraining, post-training targets task- and embodiment-specific adaptation, making it particularly sensitive to data quality. In practice, collected robot datasets often contain heterogeneous errors, including execution mistakes, sensor drift, and timestamp misalignment, which can impair post-training and policy performance. Manual inspection is costly, while existing data-cleaning methods are typically tailored to particular corruption types. To address these challenges, we introduce \textsc{RoboDrop}, a data-curation framework that audits supervision using local gradient compatibility measured along the training trajectory as a proxy for its effect on post-training performance. During a one-epoch warm-up run, RoboDrop scores each candidate sample online by comparing its gradient with those of task-semantic and visually matched validation samples. The resulting sample scores are aggregated at the episode level, and a simple automatic post-processing rule converts them into filtering decisions. We evaluate RoboDrop on controlled observation--action corruptions, naturally suboptimal demonstrations in simulation, and real-robot datasets containing non-expert collection errors. Across these settings, RoboDrop more accurately distinguishes unreliable demonstrations than prior methods, while post-training on the curated data consistently yields stronger downstream policies, with average real-robot rollout success rising from $35.0\%$ to $67.5\%$. These results establish training-trajectory-aware, context-conditioned supervision auditing as an effective approach to robust VLA post-training.

发表机构

  • Tsinghua University(清华大学)
  • Striding AI(行云智能)

机构由 AI 辅助整理,请以论文原文为准。

↑