发表机构
Oregon State University(俄勒冈州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
PhysMamba首次将选择性状态空间模型用于学习式铰接体模拟,无需速度输入即可预测全身状态,在部分观测下精度高且速度快,可作为可微物理模块集成于视频网格恢复。
AI 中文摘要
我们首次引入了一种基于选择性状态空间模型(SSM)的学习式铰接体模拟器,称为PhysMamba。PhysMamba根据位置、旋转和关节动作历史预测下一帧全身状态,无需速度输入。我们在部分观测和全观测输入下比较了四种架构,并采用了三种训练协议。从头开始的滚动训练协议使Mamba2在部分观测下具有强大的短期和中程精度(s10 = 43 mm,2/50发散),而两阶段基于教师的滚动训练协议稳定了GRU,但对Mamba2失败。借助CUDA图编译,Mamba2在H100 GPU上达到每帧0.107毫秒(9,334 FPS),在GRU未编译吞吐量的1.1倍以内,为30 Hz HMR流程增加了不到1%的延迟,并使其能够作为可微物理模块集成到基于视频的网格恢复中。
英文摘要
We introduce the first learned articulated body simulator based on a selective state space model (SSM), called PhysMamba. PhysMamba predicts next-frame full-body state from position, rotation, and joint-action history, without velocity inputs. We compare four architectures under partial- and full-observation inputs and three training protocols. The from-scratch rollout training protocol gives Mamba2 strong short- and mid-horizon accuracy under partial observation (s10 = 43 mm, 2/50 diverged), while the two-stage teacher-based rollout protocol stabilizes GRU but fails for Mamba2. With CUDA graph compilation, Mamba2 reaches 0.107 ms per frame (9,334 FPS) on an H100 GPU, within 1.1$\times$ of GRU's un-compiled throughput, adding under 1% latency to a 30 Hz HMR pipeline and enabling integration as a differentiable physics module for video-based mesh recovery.
CommentsCVPR 2026 Workshop on Physically Grounded Human Perception and Modeling -- 1st PhysHuman