发表机构
Seoul National University(首尔大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对需多关节状态运动观测的现有关节对象重建方法的局限,提出从单一静止闭合配置重建关节对象的方法,通过显式网格融合多模型输出,结合视频扩散模型实现无运动观测下的关节参数估计,性能与多种基线相当。
AI 中文摘要
构建交互式数字孪生需要同时恢复3D几何和控制对象关节运动的运动学结构。然而现有关节对象重建方法需要从多个关节状态中显式观测到运动。我们提出一种静止状态公式,仅从单一闭合配置重建关节对象,这是一个本质上不适定的场景,需通过几何、语义和运动先验补偿运动线索的缺失。我们的框架采用显式网格作为中间表示,用于跨模型验证与融合,将视觉-语言模型和分割模型的噪声输出协调为空间一致的部件结构。为在无观测运动的情况下估计关节参数,我们使用视频扩散模型合成关节假设并通过几何一致性验证。我们的方法实现了准确的部件分解和物理上合理的关节运动,在与基于运动观测的重建方法、生成方法及模块化预训练模型基线的比较中表现具有竞争力。
英文摘要
Building interactive digital twins requires recovering both 3D geometry and the kinematic structures that govern how objects articulate. Yet existing methods for articulated object reconstruction require explicitly observable motion from multiple articulation states. We introduce a rest-state formulation that reconstructs articulated objects from a single closed configuration, an inherently ill-posed setting where geometry, semantics, and motion priors compensate for the absence of motion cues. Our framework adopts an explicit mesh as an intermediate representation for cross-model verification and fusion, reconciling noisy outputs from vision-language and segmentation models into spatially consistent part structures. To estimate joint parameters without observed motion, we use a video diffusion model to synthesize articulation hypotheses and validate them through geometric consistency. Our approach achieves accurate part decomposition and physically plausible articulation, performing competitively with motion-observing reconstruction-based, generation-based, and modular pretrained-model baselines.
CommentsECCV 2026