发表机构
Alaya Lab; Shanghai Innovation Institute(Alaya实验室; 上海创新研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ShadowDancer通过阴影对与跨阴影预测学习统一动力学表示,实现视频世界模型的任意动作帧级控制,在多动力学族上的动作迁移与长回放性能优于基线,平均盲胜率达86%。
AI 中文摘要
我们提出ShadowDancer,一种用于交互式视频世界模型的任意动作、帧级控制的新方法。现有方法存在表示层面的障碍:现有接口要么对动作的编码较为松散,将动作的展开方式交由模型自行调整;要么通过结构化信号精确编码动作,但这类信号仅适用于特定动作族且难以获取,因此在不同动力学之间实现精确控制仍不切实际。演示视频是解决该问题的自然方案,它能逐帧指定任意动力学;然而视频仅能通过特定外观展现其动力学,是底层动力学的单一“阴影”,因此从演示中学习的动作难以迁移到新场景。ShadowDancer通过两项关键创新解决这一问题:(1)阴影对(shadow pairs),即同一动力学在独立重采样外观下的视频对,由我们的Shadow Library大规模构建,当某一动力学族能构建此类对时,即可对其实现精确控制;(2)跨阴影预测,通过从一个阴影预测另一个阴影来学习动作,这样重采样得到的内容会被按构造丢弃,而保留的内容则成为动作,从而得到驱动块因果世界模型的统一动力学表示。因此,任何演示片段都可成为可复用的动作资产,无需动作标签、运动估计器或微调即可在新环境中重放。实验表明,在不同动力学族上,相较于强大的潜在动作和交互式世界模型基线,ShadowDancer在动作迁移和长动作回放上表现更优,在回放比较中平均盲胜率达86%。我们在此httpsURL处展示视频结果。
英文摘要
We present ShadowDancer, a novel approach to any-action, frame-level control of interactive video world models. The obstacle is representational: existing interfaces either encode an action loosely, leaving how it unfolds for the model to improvise, or encode it exactly through structured signals that serve one family and are hard to acquire, so precise control across diverse dynamics remains impractical. Demonstration videos are the natural remedy, specifying any dynamics frame by frame; yet a video shows its dynamics only through one particular appearance, a single shadow of the underlying dynamics, so actions learned from demonstrations transfer poorly to new scenes. ShadowDancer addresses this with two key innovations: (1) shadow pairs, video pairs that replay the same dynamics under independently resampled appearance, constructed at scale by our Shadow Library, so that a dynamics family becomes controllable exactly when such pairs can be constructed for it; and (2) cross-shadow prediction, which learns actions by predicting one shadow from the other, so that whatever the pairing resamples is discarded by construction and whatever it preserves becomes the action, yielding a unified dynamics representation that drives a block-causal world model. Any demonstrated clip thus becomes a reusable action asset, replayed in new environments without action labels, motion estimators, or fine-tuning. Experiments demonstrate improved action transfer and long action rollout over strong latent-action and interactive world model baselines across diverse dynamics families, with an average blinded win rate of 86% in rollout comparisons. We show video results at https://ShadowDancer-1.github.io
Commentshttps://ShadowDancer-1.github.io