发表机构
Tongyi Lab(通义实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出用于上身人体-物体交互的实时以人为中心世界模型,采用连续-离散联合控制方案,结合多尺度运动编码的连续人体状态控制与语言编码离散交互状态控制,能精确建模人物与环境交互,提炼后实现高效实时推理,实验验证了其优势。
AI 中文摘要
我们提出了一种用于上身交互生成的实时以人为中心的世界模型,旨在合成以人物为中心的连贯局部世界动态,其中身体、手部和面部动作与可控的人体-物体离散交互共同演化。为此,我们采用了一种连续-离散联合控制方案,包括连续人体状态和离散交互状态两个互补组件。对于连续人体状态控制,引入基于多尺度运动编码的统一隐式表示,融合上半身、手部和面部的运动潜变量。对于离散物体交互状态控制,用少量语言编码的离散交互状态表示物体接触,文本作为显式交互状态命令,还构建了专用渲染管道来监督离散交互状态。通过结合两者,模型能精确建模人物与局部环境的移动和交互,包括附近场景状态的可控变化。最后对模型进行提炼以实现高效实时推理,在两个H100 GPU上达到25 FPS。实验证明了模型在细粒度运动保真度、手部-物体协调和实时交互方面的改进,朝着实时以人为中心的世界建模迈出了实际一步。
英文摘要
We present a real-time human-centric world model for upper-body interactive generation, aiming to synthesize coherent local world dynamics centered on a person, where coordinated body, hand, and facial motions evolve jointly with controllable human-object discrete interaction. To this end, we adopt a continuous-discrete joint control scheme with two complementary components: a continuous human state and a discrete interaction state. For continuous human-state control, we introduce a unified implicit representation based on multi-scale motion encoding, in which motion latents from the upper body, hands, and face are fused into a shared latent space. This multi-scale design improves expressiveness across different spatial scales, captures fine-grained human dynamics more effectively, and enables direct control without explicit retargeting. For discrete object interaction-state control, we represent object contact using a small set of language-encoded discrete interaction states, where text serves as an explicit interaction-state command, such as \emph{no contact} or \emph{grasp}, rather than an open-ended generation prompt, and we further construct a dedicated rendering pipeline for human-object interaction data to supervise such discrete interaction states. By combining continuous implicit human-state control with discrete interaction-state control, our model enables precise modeling of how a person moves and interacts with the local environment, including controllable changes to nearby scene states. Finally, we distill the model for efficient streaming real-time inference, achieving 25 FPS on two H100 GPUs. Experiments demonstrate improved fine-grained motion fidelity, more realistic hand-object coordination, and effective real-time interaction, establishing a practical step beyond motion reproduction toward real-time human-centric world modeling.