AI 中文总结
ViP-Rig是支持先提示后绑定与结果引导编辑的视觉提示框架,通过两阶段生成实现精准可控绑定,在数据集上表现优于几何基线,具备显式局部控制能力。
AI 中文摘要
绑定(Rigging)本质上依赖于任务,因为同一网格在不同动画任务中可能需要不同的骨骼结构和变形行为。实际操作中,艺术家通常会检查初始绑定,反复编辑其骨骼结构与变形行为以满足特定动画需求。现有自动方法主要从几何生成合理的绑定,对生成的骨骼和变形行为的显式控制有限。本研究提出ViP-Rig,这是一个视觉提示框架,支持先提示后绑定以及结果引导编辑,方法是将从用户绘制或编辑的2D骨骼和刚性提示中提取的特征注入冻结的预训练骨干网络。具体而言,ViP-Rig包含两个阶段:骨骼生成与蒙皮预测。第一阶段,骨骼草图通过密到疏视觉提示编码处理,生成紧凑的固定长度条件令牌;这些令牌通过门控适配器注入冻结的预训练自回归生成器,以控制关节位置和分支结构,同时保留生成器的几何先验。第二阶段,刚性图采用相同的视觉编码设计处理,而预训练蒙皮骨干保持冻结,生成的令牌对称注入点和关节流,以调节点-关节兼容性及蒙皮权重。在Articulation-XL2.0上的实验以及在ModelsResource上的零样本评估显示,在提示引导评估下,ViP-Rig比基于几何的基线更准确地恢复目标骨骼和蒙皮权重,定性结果进一步证明其在先提示后绑定和结果引导编辑中均具备显式且局部的控制能力。
英文摘要
Rigging is inherently task-dependent because the same mesh may require different skeletons and deformation behaviors across animation tasks. In practice, artists often inspect an initial rig and repeatedly edit its skeletal structure and deformation behavior to meet specific animation requirements. Existing automatic methods primarily generate a plausible rig from geometry, offering limited explicit control over the resulting skeleton and deformation behavior. In this work, we present ViP-Rig, a visual-prompted framework that supports both prompt-first rigging and result-guided editing by injecting features extracted from user-drawn or edited 2D skeletal and rigidity prompts into frozen pretrained backbones. Specifically, ViP-Rig consists of two stages, Skeleton Generation and Skinning Prediction. In the first stage, the skeletal sketch is processed by the Dense-to-Compact Visual Prompt Encoding to produce compact, fixed-length conditioning tokens. The resulting tokens are injected into a frozen pretrained autoregressive generator through gated adapters to control joint placement and branching structure while preserving the generator's geometric prior. In the second stage, the rigidity map is processed using the same visual encoding design, while the pretrained skinning backbone remains frozen. The resulting tokens are symmetrically injected into the point and joint streams to modulate point-joint compatibility and the resulting skinning weights. Experiments on Articulation-XL2.0 and zero-shot evaluation on ModelsResource show that ViP-Rig more accurately recovers target skeletons and skinning weights than geometry-conditioned baselines under prompt-guided evaluation. Qualitative results further demonstrate explicit and localized control in both prompt-first rigging and result-guided editing.
Comments8 pages, 4 figures. Zihan Qin and Mingze Sun contributed equally. Xianming Liu is the corresponding author