arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.02453cs.RO

模型失配下的规划验证:基于稀缺数据的可达性三难问题

Certifying Plans under Model Mismatch: A Trilemma for Reachability from Scarce Data

Yanliang Huang, Zhen Zhang, Ahmad Hafez, Wenyuan Wu, Peng Xie, Zhuoqi Zeng, Amr Alanwar

首次发表
浏览论文内容

中文总结 AI 辅助

针对仿真到现实策略的模型失配问题,本文提出ForeReach方法,推导了三难问题相关结论,可验证控制序列的可达性,在基准系统中表现优于校准基线。

中文摘要 AI 辅助

仿真到现实的策略是在标称动力学下设计的,但目标系统的试验可能仅产生少量孤立的单步转移。我们研究固定控制序列的执行前验证,例如由学习策略生成的动作块。如果该序列到达未观测到的状态输入区域,观测结果将与沿该区域轨迹任意大程度分离的目标系统保持一致。任何对所有目标系统都可靠的确定性验证器,要么必须拒绝验证,要么返回具有任意大投影宽度的可达管。对于目标-标称模型误差的有界光滑类,我们推导了依赖于规划的有限投影宽度下界。这些结果揭示了均匀轨迹包含、有限投影宽度以及观测之外无约束模型误差行为之间的三难问题。ForeReach需要提供模型误差的逐分量 Lipschitz 界。观测到的转移对可以反驳该声明,但无法在观测位置之外确立它。在声明有效的条件下,我们的方法构建模型误差的集合成员包络,传播 zonotopic 可达管,且仅当传播仍在验证域内且每个投影管切片避开不安全集时才进行验证。在两个基准系统中,校准基线在数据支持之外失去轨迹包含后可能仍保持较窄,而我们的方法拒绝验证无支持序列,并在相关目标数据和足够的障碍物间隙可用时恢复验证。

英文摘要

Sim-to-real policies are designed under nominal dynamics, but target-system trials may yield only a few isolated one-step transitions. We study pre-execution certification of a fixed control sequence, such as an action chunk produced by a learned policy. If the sequence reaches an unobserved state-input region, the observations remain consistent with target systems whose trajectories separate along it by an arbitrarily large amount. Any deterministic certifier sound for all of them must then decline to certify or return a reachable tube with arbitrarily large projected width. For bounded smooth classes of the target-nominal model error, we derive a finite plan-dependent projected-width lower bound. These results expose a trilemma among uniform trajectory containment, finite projected width, and unrestricted model-error behavior beyond the observations. ForeReach requires a supplied componentwise Lipschitz bound on the model error. Observed transition pairs can refute this declaration but cannot establish it outside the observed locations. Conditional on a valid declaration, our method constructs a set-membership envelope for the model error, propagates a zonotopic reachable tube, and certifies only when propagation remains within the certification domain and every projected tube slice avoids the unsafe set. In two benchmark systems, calibration baselines may remain narrow after losing trajectory containment outside data support, whereas our method declines to certify unsupported sequences and recovers certification when relevant target data and sufficient obstacle clearance are available.

↑