STATERA:通过冻结时间管状块实现零样本仿真到真实运动学隐藏质量估计
STATERA: Hidden Mass Estimation via Zero-Shot Sim-to-Real Kinematics using Frozen Temporal Tubelets
浏览论文内容
中文总结 AI 辅助
提出STATERA,利用冻结的V-JEPA骨干和时间管状块,在零样本仿真到真实场景中从单目视频估计隐藏质心,显著提升物理捕获率。
中文摘要 AI 辅助
针对帧级外观预训练的视觉模型往往难以从运动中推断隐藏物理属性的问题,我们研究了在短单目视频中对不透明、非对称刚体进行质心定位,其中在自遮挡情况下表面线索和点跟踪不可靠。我们提出STATERA,该方法采用预训练视频骨干网络(V-JEPA),保持大部分权重冻结,并配备轻量级时间管状块混合器,以预测逐帧质心热图和轨迹。为支持该任务,我们引入了HiddenMass基准,包含50K条MuJoCo轨迹和63个序列的真实世界测试集,并带有物理校准的质心真值。在仿真中,STATERA-50K-Sigma将归一化质心误差从41.7%(DINOv2)降至25.2%。在零样本仿真到真实迁移中,我们观察到监督中的基本权衡:相位感知目标可诱发双峰预测,而相位无关目标可能坍缩至统计上安全的质心。尽管如此,我们相位感知的STATERA-50K-Crescent是唯一在所有评估方法中持续向真实隐藏偏移移动的方法。虽然这导致单目向量过冲伪影,使绝对欧氏误差相比静态几何质心略有增加,但将物理捕获率从2.6%提升至41.0%。这些结果表明,冻结的时间表示能更好地将惯性动力学与视觉几何分离,用于隐藏参数估计。
英文摘要
Vision models pretrained for frame-level appearance often struggle to infer hidden physical properties from motion. We study center-of-mass (CoM) localization for opaque, asymmetric rigid bodies from short monocular videos, where surface cues and point tracking are unreliable under self-occlusion. We propose STATERA, which adapts a pretrained video backbone (V-JEPA) with mostly frozen weights and a lightweight temporal tubelet mixer to predict per-frame CoM heatmaps and trajectories. To support this task, we introduce the HiddenMass Benchmark, comprising 50K MuJoCo trajectories and a 63-sequence real-world test set with physically calibrated CoM ground truth. In simulation, STATERA-50K-Sigma improves normalized CoM error from 41.7% (DINOv2) to 25.2%. In zero-shot sim-to-real transfer, we observe a fundamental trade-off in supervision: phase-aware targets can induce bimodal predictions, while phase-agnostic targets can collapse toward statistically safe centroids. Nevertheless, our phase-aware STATERA-50K-Crescent is the only evaluated method that demonstrates consistent movement toward the true hidden offset. While this leads to a monocular vector overshoot artifact that marginally increases absolute Euclidean error compared to a static geometric centroid, it improves physics capture from 2.6% to 41.0%. These results suggest that frozen temporal representations can better separate inertial dynamics from visual geometry for hidden-parameter estimation.