AI 中文总结
针对VLA模型易受在线干扰的问题,提出无需训练的CoRe框架,通过反事实想象与最小重对齐实现无试错恢复,大幅提升任务成功率并减少物理恢复操作。
AI 中文摘要
视觉-语言-动作(VLA)模型提升了机器人操作的灵活性与通用性,但对在线干扰仍较为脆弱,例如任务目标、场景配置或机器人状态的变化。现有恢复方法通常需要故障数据、策略重训练或外部纠正智能体,带来额外的数据需求与执行风险。我们提出反事实重对齐(Counterfactual Realignment, CoRe),这是一种无需训练的框架,可在推理时恢复冻结的VLA模型,且无需故障数据。检测到偏差后,CoRe会利用合成观测替代物理执行,从最近的可行状态出发,模拟策略朝向当前目标的后续走向,随后最小程度地调整机器人与场景,以重新接入该想象的后续流程,再将控制权交还给策略。因此,恢复过程无需物理试错,可保留已完成的任务进度,且能统一处理 episode 中途的指令变更与物理扰动。在多个模拟器、VLA骨干网络及真实场景中开展的大量实验表明,CoRe可将成功率提升最多85.0个百分点,接近名义水平,同时减少42.2%的物理恢复操作,且无需策略微调或针对特定故障的恢复训练。
英文摘要
Vision-language-action (VLA) models have improved the flexibility and generality of robotic manipulation, yet they remain fragile to online disruptions, such as changes in task goal, scene configuration, or robot state. Existing recovery methods often require failure data, policy retraining, or external corrective agents, introducing additional data requirements and execution risks. We propose Counterfactual Realignment (CoRe), a training-free framework that recovers a frozen VLA at inference time without failure data. Upon detecting a deviation, CoRe imagines how the policy would continue toward the current goal from a recent viable state, using synthesized observations in place of physical execution, and then minimally realigns the robot and scene to rejoin this imagined continuation before returning control to the policy. Recovery is therefore planned without physical trial-and-error, preserves completed task progress, and handles both mid-episode instruction changes and physical perturbations in a unified manner. Extensive experiments across multiple simulators, VLA backbones, and real-world settings show that CoRe improves success rates by up to 85.0 percentage points to near-nominal levels while reducing physical restorations by 42.2%, without policy fine-tuning or failure-specific recovery training.