发表机构
California Institute of Technology(加州理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出PAC-MAN感知感知CBF-RL框架,结合控制屏障安全与机载传感,在人形机器人躲避球任务中,基于Unitree G1实现95%投掷成功率,验证了感知可观测性对屏障结构性能的影响。
AI 中文摘要
我们提出了PAC-MAN,这是一种感知感知CBF-RL框架,将控制屏障安全与部署时真实的机载传感相结合,用于人形机器人躲避球的全身控制。所部署的策略仅将球视为头戴式摄像机的分割掩码深度,而训练时的CBF指导表示与每个身体连杆的间隙,对抗性运动先验对生成的规避反射进行正则化。我们在受控的任意连杆接触基准上评估,该基准带有两种模式的种子投掷:单次投掷和部署循环,其中机器人返回其站位并在投掷之间恢复。在该基准上,该策略与特权状态预言机仅差几分:仅固定机载摄像机就足以实现规避。我们发现,可用的屏障结构取决于感知可观测性:Joint-CBF在球状态准确时提供最佳性能,当仅用作训练指导时在固定摄像机观测下性能下降,而使用球跟踪云台或特权运行时过滤器可恢复性能。因此,我们在现实世界中零样本部署了轻量Link-CBF策略在Unitree G1上,该策略可容忍不完美感知,在95%的投掷中成功,并使用语义分割躲避不同的球。
英文摘要
We present PAC-MAN, a perception-aware CBF-RL framework that couples control-barrier safety with deployment-realistic onboard sensing for whole-body humanoid dodgeball. The deployed policy sees the ball only as segmentation-masked depth from a head-mounted camera, while training-time CBF guidance represents clearance to every body link, and an adversarial motion prior regularizes the resulting evasive reflexes. We evaluate on a controlled any-link contact benchmark with seeded throws in two regimes: single throws and a deployment loop in which the robot walks back to its station and recovers between throws. On this benchmark, the policy comes within a few points of a privileged state oracle: a fixed onboard camera alone is adequate for evasion. We find that usable barrier structure depends on perceptual observability: Joint-CBF gives the best performance with accurate ball states, degrades under fixed-camera observations when used only as training guidance, and recovers with a ball-tracking gimbal or privileged runtime filter. We therefore deploy a lightweight Link-CBF policy zero-shot on the Unitree G1 in the real world, where it tolerates imperfect perception, succeeds on 95% of throws, and uses semantic segmentation to dodge different balls.
CommentsWebsite at https://lzyang2000.github.io/perceptive_cbf_rl/