三维人体运动生成的流式多轨时间线控制
Streaming Multi-Track Timeline Control for 3D Human Motion Generation
浏览论文内容
中文总结 AI 辅助
针对交互式流式指令下的人体运动生成,提出TimelineControl方法,通过间隔感知条件与部分感知去噪实现多轨时间线控制,并构建TimelineMotion数据集,实验验证了其语义对齐与时间遵循性的提升。
中文摘要 AI 辅助
文本驱动的人体运动生成已取得显著进展,但大多数方法假设指令在合成之前可用。交互式应用要求在对正在进行的动作继续响应的同时响应新指令,例如在走路时接电话。现有方法处理流式生成或同时组合,但未明确将流式指令到达与独立计时的重叠动作相结合。我们引入流式多轨时间线控制,并提出TimelineControl以将新指令与正在进行的动作相结合。间隔感知条件保持指令时序,而因果部分结构化表示和部分感知去噪协调跨身体区域的并发动作。我们还构建了TimelineMotion,一个具有重叠指令间隔和身体部位注释的数据集。在TimelineMotion和MTT上的实验表明,与评估的流式基线(包括在同一数据上重新训练的模型)相比,语义对齐和时间遵循性有所改善。消融研究和人工评估验证了我们的设计,并辅以空间条件和人形执行演示。我们的代码、数据和模型将公开可用。
英文摘要
Text-driven human motion generation has advanced substantially, yet most methods assume instructions are available before synthesis. Interactive applications require responding to new instructions while continuing ongoing actions, such as answering a phone while walking. Existing approaches address streaming generation or simultaneous composition without explicitly combining streaming instruction arrival with independently timed, overlapping actions. We introduce streaming multi-track timeline control and propose TimelineControl to incorporate new instructions alongside ongoing actions. Interval-aware conditioning preserves instruction timing, while causal part-structured representations and part-aware denoising coordinate concurrent actions across body regions. We also construct TimelineMotion, a dataset with overlapping instruction intervals and body-part annotations. Experiments on TimelineMotion and MTT demonstrate improved semantic alignment and temporal adherence over evaluated streaming baselines, including models retrained on the same data. Ablations and human evaluations validate our design, complemented by spatial conditioning and humanoid execution demonstrations. Our code, data and models will become publicly available.
发表机构
- Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)(穆罕默德·本·扎耶德人工智能大学)
- LIGM, École des Ponts, IP Paris, Univ Gustave Eiffel, CNRS(LIGM,巴黎高科桥梁学院,巴黎理工学院,古斯塔夫·埃菲尔大学,法国国家科学研究中心)
机构由 AI 辅助整理,请以论文原文为准。