Flex-$π$:具备计算灵活性的多流世界-动作模型
Flex-$π$: A Multi-Stream World-Action Model with Compute Flexibility
浏览论文内容
中文总结 AI 辅助
该研究提出Flex-$π$多流世界-动作模型,利用冻结的视频生成VAE实现多模态监督,通过混合专家Transformer与每流丢弃机制提升效率,在双臂操作任务上性能优于基线模型且运行更快。
中文摘要 AI 辅助
世界-动作模型(World-action models, WAMs)通过预测未来来优化决策,但几乎所有此类模型仅预测RGB潜变量,且仅针对像素重建进行训练,未明确考虑3D几何或对象语义操作的需求。我们发现了一个令人惊喜的意外收获:经冻结的视频生成VAE(变分自编码器)在未进行点云图特定训练的情况下,几乎能无损地对3D点云图进行编码,同时也可对RGB进行编码。这使得我们能够在无需新增传感器、预训练或推理延迟的情况下,在3D几何、以对象为中心的DINO语义以及RGB上对参数规模达60亿的WAM模型Flex-$π$进行监督训练。所有视觉信号均被投影至该共享潜空间,并在混合专家Transformer(Mixture-of-Transformers)主干网络中与动作一同进行去噪处理;通过跨模态强制的每流丢弃机制,单个训练好的检查点可在任意流子集上运行,从仅动作的快速模式到完整的联合生成都可实现。最终得到的策略具有极高的演示效率和良好的泛化能力,在分布内和分布外的灵巧、精确的真实世界双臂操作任务上,击败了最强基线模型,性能提升达2至7倍,同时运行速度快于$π_{0.5}$。项目网站:this https URL
英文摘要
World-action models (WAMs) predict the future to act better, but nearly all of them predict only RGB latents, trained purely for pixel reconstruction, with no explicit signal for the 3D geometry or object semantics manipulation needs. We find a surprising free lunch: the same frozen video-generation VAE that encodes RGB also encodes 3D pointmaps almost losslessly, with no pointmap-specific training at all. This lets us supervise Flex-$π$, a 6B-parameter WAM, on 3D geometry and object-centric DINO semantics alongside RGB, at no cost in new sensors, new pre-training, or inference latency. Every visual signal is projected into this shared latent space and denoised jointly with actions inside a Mixture-of-Transformers backbone; per-stream dropout with cross-modality forcing then lets a single trained checkpoint run on any subset of these streams, from a fast action-only mode to full joint generation. The result is a policy that is exceptionally demonstration-efficient and generalizes well, beating the strongest baselines by up to 2-7$\times$ on dexterous, precise, real-world bimanual manipulation tasks both in and out of distribution, all while running faster than $π_{0.5}$. Our project website: https://flex-pi.github.io/
发表机构
- University of Washington(华盛顿大学)
- Allen Institute for AI(艾伦人工智能研究所)
机构由 AI 辅助整理,请以论文原文为准。