发表机构
Huawei Heisenberg Research Center; TU Berlin; UCL Center for AI(华为海森堡研究中心; 柏林工业大学; 伦敦大学学院人工智能中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对VLA策略中安全动作不可行的问题,提出可行性-似然差距概念,开发警报触发、无需训练的VICS-G重新排序器,在Safety-CHORES上降低安全成本1.9%-57.5%且性能损失极小。
AI 中文摘要
安全动作未必是可行的动作。在冻结的视觉-语言-动作(VLA)策略下,一个动作可能具有高概率且局部可接受,但留下没有策略支持的安全任务完成路径。我们称此为可行性-似然差距:似然度对当前动作进行排序,而可行性取决于动作之后保留的未来。我们推导了历史条件策略-环境轨迹定律在限制为安全任务完成时的精确下一块边缘。推导揭示了一个依赖于候选的可行未来质量,具有两个作用:其支撑记录了在冻结延续过程下安全完成是否仍然可能,其大小衡量了保留的加权安全完成质量有多少。在线精确评估不切实际,因此我们开发了一种选择性有限候选近似,推导了恢复最佳保留可行候选的条件,并将其实现为警报触发、无需训练的重新排序器。在Safety-CHORES上,VICS-G在六种设置中将平均累积安全成本降低了1.9%-57.5%,同时在成功率上保持在策略采样的2.5个百分点以内,在平均情节长度上保持在0.82步以内。所得解码器与精确的策略相对安全完成目标相关联,但既不需要对基础策略进行重新训练,也不需要在线轨迹回放。
英文摘要
A safe action is not necessarily a viable one. A frozen vision-language-action (VLA) policy can favor a locally admissible move that leaves no policy-supported route to safe task completion. We call this the feasibility-likelihood gap: likelihood ranks the next move, while feasibility depends on the futures it leaves open. To bring those futures into the decision, we derive the exact next-block marginal of the history-conditioned policy-environment trajectory law restricted to safe task completion. The derivation reveals a candidate-dependent feasible-future mass: its support records whether safe completion remains possible under the frozen continuation process, while its magnitude measures how much weighted safe-completion mass remains. Since exact evaluation is impractical online, we develop a selective finite-candidate approximation and establish conditions for recovering the best retained viable candidate. Our alarm-triggered, training-free reranker VICS-G lowers mean cumulative safety cost by 1.9%-57.5% across six Safety-CHORES settings while remaining within 2.5 percentage points of policy sampling in success and 0.82 steps in mean episode length. Our approach offers a promising and practical path toward safer task completion, grounded in an exact policy-relative target yet requiring neither policy retraining nor online rollouts.