发表机构
Harvard University; Independent Researcher(哈佛大学; 独立研究者)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究交互式世界模型精确状态服务中的调度问题,提出 WorldMove 方法,能快速迁移缓存且保证位相同,通过可接受性条件等解决调度难题,实现跨传输和验证的联合调度,提升集群可调度性。
AI 中文摘要
持久的交互式世界模型将其运行状态驻留在为其服务的 GPU 上,即一个多千兆字节的注意力缓存,几乎在每个生成步骤都会被重写。该状态无法在交互时间内重新计算,也不能在不改变世界的情况下近似计算,因此实时会话会固定其设备。固定是一个调度问题。WorldMove 在一种保证下移动实时会话:目标与源位相同,否则不安装任何内容。它在同节点内 18.8 毫秒内重新定位缓存,比保存/加载快 101 倍。在 100 Gb 网络上保持经校验和验证的 92.1 - 94.8 Gb/s 的速率。以该速率,缓存适合在一个交互块内。迁移一个正在主动生成的会话时,它在块边界收敛,目标逐位继续世界。一个可接受性条件决定每次移动。移动必须在读取范围内完成,通过覆盖状态及其脏速率的带宽。提升到集群可调度性测试,它控制一个整合循环,在两个提供商之间执行了 48 次迁移中的 48 次位相同迁移。有两个结构约束。位精确性仅在一个 GPU 架构的受控配置内存在,因此移动状态是在交互时间内精确保存它唯一的方法。在此网络上,验证不能隐藏在线路中。接收路径校验和在扇入情况下会在协议时间尺度上使传输停顿,未调度的集中广播会在每个传输字节保持正确的情况下悄悄使接收器崩溃。一个感知集中广播的准入控制器在提供负载的 0 到 1.4 倍时保持零丢失,并将过载作为拒绝丢弃。无损 GPU 编解码器拓宽了准入门,使其适用于原始运动无法使用的网络。我们分别对服务循环和移动器进行端到端测试。它们在一个网络上的组合尚未构建。精确状态弹性是一个跨传输和验证的联合调度问题。
英文摘要
A persistent interactive world model keeps its running state resident on the GPU that serves it: a multi-gigabyte attention cache, almost all of it rewritten at every generation step. That state cannot be recomputed in interactive time or approximated without changing the world, so a live session pins its device. The pin is a scheduling problem. WorldMove moves a live session under one guarantee: the destination is bit-identical to the source, or nothing is installed. It relocates the cache in 18.8 ms same-node, 101x faster than save/load. It holds a checksum-verified 92.1-94.8 Gb/s on a 100 Gb fabric. At that rate the cache fits inside one interactive block. Migrating an actively generating session, it converges at a block boundary and the destination continues the world bit for bit. An admissibility condition decides each move. The move must complete inside the readout horizon, over bandwidth that covers the state plus its dirty rate. Lifted to a fleet schedulability test, it governed a consolidation loop that executed 48 of 48 migrations bit-identical across two providers. Two constraints are structural. Bit-exactness survives only inside a controlled configuration of one GPU architecture, so moving the state is the only way to preserve it exactly in interactive time. Verification cannot hide inside the wire on this fabric. Receive-path checksums stall the transport at protocol timescales under fan-in, and unscheduled incast silently collapses a receiver while every delivered byte stays correct. An incast-aware admission controller holds zero misses to 1.4x offered load and sheds overload as rejects. A lossless GPU codec widens the admission gate to fabrics raw motion cannot use. We exercise the serving loop and the mover separately, each end to end. Their composition on one fabric is unbuilt. Exact-state elasticity is a joint scheduling problem over transport and verification.
Comments20 pages. Extended version