AI 中文总结
ArtiMo是一种智能体驱动的零样本框架,结合URDF运动学约束与LLMs/VLMs,生成文本引导的关节网格动画,通过视觉自改进机制修正错误,在新基准数据集上优于基线。
AI 中文摘要
通过文本生成关节型三维网格动画需满足严格的运动学约束、建模各部件间的因果交互并实现指令保真度。由于缺乏该任务的专用训练数据和显式关节监督,现有的数据驱动网格动画方法大多不适用于该场景。为解决此问题,我们提出ArtiMo,一种用于文本引导关节网格动画的新型智能体驱动框架。ArtiMo以零样本方式运行,开发了由大语言模型(LLMs)和视觉语言模型(VLMs)驱动的智能体流程来协调运动生成。通过将URDF的显式运动学约束与智能体的推理和规划能力相结合,它能有效生成具有因果一致性的部件运动和交互,无需模型微调。为确保运动正确性,智能体还利用了视觉自改进机制:将生成的动画渲染为紧凑的关键帧和运动线索,使VLM能够迭代诊断和修正错误。此外,我们贡献了一个包含21个关节物体类别的新基准数据集,具有丰富因果关系的高质量运动标注。大量实验表明,ArtiMo显著优于基线方法,尤其在复杂的因果驱动运动上表现突出。项目页面可通过此https URL访问。
英文摘要
Animating articulated 3D meshes via text requires satisfying strict kinematic constraints, modeling causal interactions between parts, and achieving instruction fidelity. Due to the absence of task-specific training data and explicit articulation supervision, existing data-driven mesh animation methods are largely inapplicable to this setting. To address this, we propose ArtiMo, a novel agent-driven framework for text-guided articulated mesh animation. Operating in a zero-shot manner, ArtiMo develops an agentic pipeline powered by Large Language and Vision-Language Models (LLMs/VLMs) to orchestrate motion generation. By synergizing the explicit kinematic constraints of URDF with the agent's reasoning and planning capabilities, it effectively produces causally coherent part motions and interactions without requiring model fine-tuning. To ensure motion correctness, the agent additionally utilizes a visual self-improvement mechanism: generated animations are rendered into compact keyframes and motion cues, enabling the VLM to iteratively diagnose and correct errors. Furthermore, we contribute a new benchmark dataset spanning 21 articulated object categories, featuring high-quality motion annotations enriched with causal relationships. Extensive experiments demonstrate that ArtiMo significantly outperforms baselines, particularly on complex, causally driven motions. The project page is available at https://zou-2004.github.io/ArtiMo/.