arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

缓存感知的Conv3D降级:跨嵌入式世界模型解码器

Cache-Aware Conv3D Lowering Across Embedded World-Model Decoders

Jiaming Zhang, Wu Yang, Shuai Tao, Wulong Liu

arXiv 2609.31938首次发表:更新:

AI 中文总结

提出缓存感知的Conv3D降级为批处理Conv2D,在Jetson AGX Orin上加速VAE解码约7倍,降低生成延迟超2倍,无需修改模型。

AI 中文摘要

生成式世界模型可以为具身规划提供视觉回放,但它们在边缘设备上的可行性不仅取决于学习到的模型,还取决于执行运行时如何表示其操作。我们引入了一种缓存感知的降级方法,将支持的因果Conv3D调用表示为批处理的空间Conv2D操作,同时保留预训练权重、时间缓存语义、卷积参数、偏置放置和输出布局。在64GB NVIDIA Jetson AGX Orin上完整的Cosmos3-Edge图像到视频流程中,所提出的路线将VAE解码加速约7倍,并将完整生成延迟降低超过2倍,而重复的解码器评估保持完整的快速路径覆盖,无需回退。未改变的降级方法还改善了Cosmos3-Nano,并迁移到LingBot-World架构上不同的Wan2.1 VAE。与完全特化的TensorRT进行的干净设备比较显示,TensorRT提供了额外的1.36倍稳态改进,但需要显著更多的逐模块和逐运行时状态的AOT特化。相同潜变量的BF16和FP32评估表征了由替代执行顺序引入的有限精度差异。总之,这些结果将缓存感知降级定位为一种轻量级运行时优化,在不修改学习模型本身的情况下,恢复了大部分可用的解码器加速。

英文摘要

Generative world models can provide visual rollouts for embodied planning, yet their feasibility on edge devices depends not only on the learned model but also on how the execution runtime represents its operations. We introduce a cache-aware lowering that expresses supported causal Conv3D calls as batched spatial Conv2D operations while preserving pretrained weights, temporal-cache semantics, convolution parameters, bias placement, and output layout. Across the complete Cosmos3-Edge image-to-video pipeline on a 64-GB NVIDIA Jetson AGX Orin, the proposed route accelerates VAE decoding by approximately $7\times$ and reduces complete-generation latency by more than $2\times$, while repeated decoder evaluations maintain complete fast-path coverage without fallbacks. The unchanged lowering also improves Cosmos3-Nano and transfers to LingBot-World's architecturally distinct Wan2.1 VAE. A clean-device comparison against fully specialized TensorRT shows that TensorRT provides a further $1.36\times$ steady-state improvement, but requires substantially greater per-module and per-runtime-state AOT specialization. Same-latent BF16 and FP32 evaluations characterize the finite-precision differences introduced by the alternative execution order. Together, these results position cache-aware lowering as a lightweight runtime optimization that recovers most of the available decoder acceleration without modifying the learned models themselves.

Comments16 pages, 1 figure, 6 tables. ECCV 2026 workshop paper

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑