发表机构
Sony Computer Science Laboratory(索尼计算机科学实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对视觉-语言-动作策略,提出六轴配对基准LIBERO-CTRL,发现复合鲁棒性无法由单轴成功率推断,需逐实例评估以揭示联合扰动下的行为变化。
AI 中文摘要
视觉-语言-动作策略通常一次只评估一种扰动,这有助于诊断其对单个分布偏移的敏感性。然而,现实世界部署可能同时涉及多种偏移,目前尚不清楚这些单独的鲁棒性测量如何组合。我们探讨是否可以从单轴评估推断复合鲁棒性。我们引入了LIBERO-CTRL,一个六轴基准,它将每个初始状态在单轴条件下与匹配的同时条件配对。这种设计揭示了两种相反的结果变化,而聚合成功率无法区分它们:突发失败(所有单轴 rollout 成功但同时 rollout 失败)和补偿成功(至少一个单轴 rollout 失败但同时 rollout 成功)。由于一种转变降低了复合成功率,而另一种提高了它,它们可能相互抵消,使得聚合复合性能看起来与单轴测量一致,即使个体结果差异显著。这两种相反的转变在聚合中可能很大程度上相互抵消:即使两种转变率之间的差异在统计上与零无显著区别,多达29.0%的匹配初始状态仍会改变结果。在六种策略和三种严重程度下,这种结果变化在最受影响的条件中达到34.5%。两种转变的相对普遍性因策略和严重程度而异,而在随机策略的独立重新评估下,转变率保持相似。因此,复合鲁棒性不能仅从聚合的单轴成功率来表征;需要匹配的逐实例评估来揭示联合扰动如何改变行为。
英文摘要
Vision-language-action policies are typically evaluated one perturbation at a time, providing a useful diagnosis of their sensitivity to individual distribution shifts. Real-world deployment, however, may involve several shifts simultaneously, and it remains unclear how these individual robustness measurements compose. We ask whether compound robustness can be inferred from single-axis evaluations. We introduce LIBERO-CTRL, a six-axis benchmark that pairs each initial state across single-axis conditions and a matched simultaneous condition. This design reveals two opposing outcome changes that aggregate success rates cannot distinguish: emergent failures, where all single-axis rollouts succeed but the simultaneous rollout fails, and compensated successes, where at least one single-axis rollout fails but the simultaneous rollout succeeds. Because one transition decreases compound success while the other increases it, they can cancel, making aggregate compound performance appear consistent with single-axis measurements even when individual outcomes differ substantially. These opposing transitions can largely cancel in aggregate: even when the difference between the two transition rates is not statistically distinguishable from zero, as many as 29.0% of matched initial states still change outcome. Across six policies and three severity levels, such outcome changes reach 34.5% in the most affected condition. The relative prevalence of the two transitions varies across policies and severities, while the transition rates remain similar under independent re-evaluation of stochastic policies. Compound robustness therefore cannot be characterized from aggregate single-axis success rates alone; matched per-instance evaluation is needed to reveal how joint perturbations alter behavior.