发表机构
University of Cambridge; Institute of Automation, Chinese Academy of Sciences; Northwestern Polytechnical University; The Hong Kong Polytechnic University(剑桥大学; 中国科学院自动化研究所; 西北工业大学; 香港理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
BendTwin是一种弯曲感知可微分弹簧-质量框架,通过引入弯曲约束提升力学稳定性,在可变形物体重建与预测任务中优于仅用轴向弹簧的PhysTwin,可用于从稀疏视角RGB-D视频构建力学保真数字孪生。
AI 中文摘要
从视频观测中重建具有力学属性的物体,可实现物理一致的动态预测,有益于机器人规划与交互。现有基于弹簧-质量的物理驱动重建方法高效且可微分,但通常仅依赖轴向弹簧,该公式过度简化了潜在结构力学,且在物理图粗化时会出现力学欠约束,限制了其保持稳定局部变形的能力。本文提出BendTwin,一种用于基于视频的可变形物体重建与未来预测的弯曲感知可微分弹簧-质量框架,该框架在局部表面三元组上引入弯曲刚度与阻尼,惩罚与静止角度的偏差并正则化高阶变形,这些弯曲约束在保留弹簧-质量系统简洁性的同时提升了力学稳定性。实验表明,BendTwin始终优于仅使用轴向弹簧的PhysTwin基准方法;消融研究进一步证明,弯曲约束在不同下采样比例下均能维持系统稳定性,并持续改进原始PhysTwin公式。总体而言,BendTwin为从稀疏视角RGB-D视频构建力学保真的数字孪生提供了有效方法。
英文摘要
Reconstructing objects with mechanical properties from video observations enables physically consistent dynamic prediction, benefiting robotics planning and interaction. Existing spring--mass based physical driven reconstruction approaches offer efficient and differentiable physical reconstruction, but they typically rely on axial springs alone. Such formulations oversimplify the underlying structural mechanics and can become mechanically under-constrained when the physical graph is coarsened, limiting their ability to preserve stable local deformation. We present BendTwin, a bending-aware differentiable spring--mass framework for video-based reconstruction and future prediction of deformable objects. BendTwin introduces bending stiffness and damping over local surface triplets, penalizing deviations from rest angles and regularizing higher-order deformation. These bending constraints improve mechanical stability while preserving the simplicity of spring--mass system. Experiments show that BendTwin consistently outperforms the axial-only PhysTwin baseline. Ablation studies further demonstrate that the bending constraints maintain system stability across different downsampling ratios and consistently improve upon the original PhysTwin formulation. Overall, BendTwin provides an effective approach for constructing mechanically faithful digital twins from sparse-view RGB-D videos.