AI 中文总结
本文提出WorldDynCache框架,通过风险估计器和提升隐式代理解决扩散世界模型推理慢的问题,在两个数据集上实现加速并取得最优生成质量。
AI 中文摘要
扩散世界模型可生成高质量的未来内容,但重复的Transformer评估会导致推理速度过慢。现有缓存方法会复用中间特征、选择性更新token,或根据局部漂移或短原生空间历史复用并外推去噪输出,这些标准会遗漏因跳过步骤累积的近似诱导隐式转移缺陷,以及隐式演化方向中依赖阶段或条件的变化。我们提出WorldDynCache,这是一个风险控制的隐式动力学近似框架,包含两个核心组件:一是轻量隐式转移风险估计器,用于跟踪近似缺陷对未来的累积影响,并根据精确锚点处观察到的反事实缺陷校准其预测;二是依赖条件和阶段的提升隐式代理,无需额外的Transformer评估即可近似隐式演化。在HunyuanVoyager-13B和Aether-5B上,WorldDynCache分别实现了4.92倍和2.15倍的加速,同时在WorldScore、PSNR、SSIM和LPIPS指标上,在对比的缓存方法中达到了最佳的生成质量。
英文摘要
Diffusion world models generate high-quality futures, but re- peated transformer evaluations make inference prohibitively slow. Existing caches reuse intermediate features, selectively update tokens, or reuse and extrapolate denoising outputs ac- cording to local drift or short native-space histories. These criteria can miss both approximation-induced latent transition defects that accumulate across skipped steps and phase- or condition-dependent changes in the direction of latent evo- lution. We propose WorldDynCache, a risk-controlled latent dynamics approximation framework with two core compo- nents. First, a lightweight latent-transition risk estimator tracks the accumulated future impact of approximation defects and calibrates its predictions against counterfactual defects ob- served at exact anchors. Second, a condition- and phase- aware lifted latent surrogate approximates latent evolution without extra transformer evaluations. On HunyuanVoyager- 13B and Aether-5B, WorldDynCache achieves 4.92 times and 2.15 times speedups, respectively, while attaining the best gen- eration quality among the compared caching methods across WorldScore, PSNR, SSIM, and LPIPS.