发表机构
Skolkovo Institute of Science and Technology(斯科尔科沃理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Kintsugi-VLA利用仿真中的状态恢复和分支,将失败轨迹转化为有针对性的恢复数据,通过干预可恢复性估计选择起始状态,显著提升恢复成功率。
AI 中文摘要
仿真通过利用特权专家生成视觉演示,无需通过手动遥操作收集每条轨迹,从而能够对视觉-语言-动作策略进行可扩展的训练。然而,此类流程通常保留成功的演示,而丢弃失败的轨迹,尽管这些失败的轨迹恰恰暴露了必须从中学习恢复的离标称状态。我们提出了Kintsugi-VLA,一个通过利用仿真中的精确状态恢复和分支,将失败的轨迹转化为有针对性的合成恢复数据的框架。对于固定的特权专家,我们将干预可恢复性定义为在模拟器恢复到给定状态后完成原始任务的概率,使用带有逐点Wilson区间的自适应蒙特卡洛延续来估计它,并刻画其沿失败轨迹的非单调演变。这些估计识别出一个观察到的终端低可恢复性前沿——即测得的可恢复性保持在阈值以下的点之后——然后用于选择信息丰富的恢复起始状态。在模拟的Franka操作任务中,在难度匹配和帧预算匹配条件下,有针对性的恢复数据分别使SmolVLA聚合恢复成功率达到34.6%和38.4%,比同一恢复窗口内的均匀采样高出5.8和6.7个百分点。在受干扰的端到端执行以及偏移的杂乱和物理条件下,观察到相同的排序,而干净任务成功率从76.8%下降到74.7%。Kintsugi-VLA展示了如何通过直接的干预测量,将失败的模拟器轨迹从被丢弃的经验转化为结构化的恢复训练数据。
英文摘要
Simulation enables scalable training of Vision-Language-Action policies by using privileged experts to generate visual demonstrations without requiring every trajectory to be collected through manual teleoperation. However, such pipelines typically retain successful demonstrations while failed rollouts are discarded, even though they expose precisely the off-nominal states from which recovery must be learned. We introduce Kintsugi-VLA, a framework for converting failed rollouts into targeted synthetic recovery data by exploiting exact state restoration and branching in simulation. For a fixed privileged expert, we define interventional recoverability as the probability of completing the original task after the simulator is restored to a given state, estimate it using adaptive Monte Carlo continuations with pointwise Wilson intervals, and characterize its non-monotonic evolution along failed trajectories. These estimates identify an observed terminal low-recoverability frontier-the point after which measured recoverability remains below a threshold-which is then used to select informative recovery starting states. In a simulated Franka manipulation task, targeted recovery data yield aggregate SmolVLA recovery success of 34.6\% and 38.4\% under difficulty- and frame-budget matching, respectively, 5.8 and 6.7 percentage points above uniform sampling within the same recovery window. The same ordering is observed under disturbed end-to-end execution and shifted clutter and physics conditions, while clean-task success decreases from 76.8\% to 74.7\%. Kintsugi-VLA demonstrates how failed simulator rollouts can be transformed from discarded experience into structured recovery-training data through direct interventional measurement.
Comments8 pages, 7 figures, 5 tables (applied on ICRA2027)