arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

异构机器人的校准预测安全性:一种带基于模型安全盾的动作条件联合嵌入预测架构(JEPA)框架

Calibrated Predictive Safety for Heterogeneous Robots: An Action-Conditioned JEPA Framework with Model-Based Safety Shields

Kaiming Zhong, Tianhua Liu, Yue Wang

arXiv 2608.17496首次发表:更新:

发表机构

Guangdong Bifang Intelligent Control Technology Co., Ltd.(广东毕方智能控制技术有限公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出带基于模型安全盾的动作条件JEPA框架,用于异构机器人,可预测候选动作的任务进展与物理风险,在仿真中提升了成功率并降低碰撞漏报率。

AI 中文摘要

视觉-语言-动作策略泛化性强,但无法提供执行时的保障;经典基于模型的规划器虽能满足运动学和几何约束,但泛化性差。本研究探讨动作条件联合嵌入预测架构(JEPA)世界模型能否在执行前预测候选动作块的任务进展与物理风险,以及将这些预测与特定本体的基于模型安全盾耦合,是否能为异构机器人生成可部署的流程。我们提出一种后退时域决策流程:(1)提议器生成K个候选动作块;(2)动作条件JEPA在冻结编码器的潜在空间中,基于本体嵌入对每个候选动作块进行回滚;(3)校准后的风险与进展头对每次回滚进行评分并报告不确定性;(4)确定性的本体专属安全盾过滤不可接受的候选;(5)回退梯处理候选集为空的情况。学习到的排序仅对可接受候选进行重新排序,执行保障来自确定性安全盾与回退梯。我们采用预注册协议在LIBERO-Long仿真环境中评估:在600回合配置下,完整框架相比仅带安全盾的基线提升了成功率,且在匹配召回率下降低了碰撞漏报率;还包含目标机器人端与边缘加速器的部署效率测量。真实机器人实验与离线重排序显著性检验仍为未来工作,详情见论文披露。

英文摘要

Vision-language-action policies generalize broadly but provide no execution-time guarantees; classical model-based planners respect kinematic and geometric constraints but generalize poorly. We study whether an action-conditioned Joint-Embedding Predictive Architecture (JEPA) world model can predict, before execution, both task progress and physical risk for candidate action chunks, and whether coupling these predictions to an embodiment-specific model-based safety shield yields a deployable pipeline for heterogeneous robots. We propose a receding-horizon decision pipeline: (1) a proposer produces K candidate action chunks; (2) an action-conditioned JEPA rolls each candidate forward in a frozen-encoder latent space conditioned on an embodiment embedding; (3) calibrated risk and progress heads score each rollout and report uncertainty; (4) a deterministic per-embodiment safety shield filters inadmissible candidates; (5) a fallback ladder handles empty-admissible-set cases. The learned ranking only reorders admissible candidates; enforcement guarantees come from the deterministic shield and fallback ladder. We evaluate with a pre-registered protocol in simulation (LIBERO-Long). In 600-episode configurations the full framework improved success over a shield-only baseline and reduced collision false negatives at matched recall. Deployment-efficiency measurements on target on-robot and edge accelerators are included. Real-robot experiments and an offline reranking significance test remain future work; see the paper for disclosures.

Comments17 pages, 9 figures. Simulation-only empirical results on LIBERO-Long (no real-robot experiments). Source, figure-generation scripts and reproducibility checklist included. Level-3 offline reranking significance test not executed; see Sec. 7 (Scope and honesty statement) for detailed disclosure

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑