通过动作单元和动作检测引导实现文本到运动的逐笔画时间控制
Per-Stroke Temporal Control for Text-to-Motion via Action Units and Action-Detection Guidance
浏览论文内容
中文总结 AI 辅助
研究文本到运动模型中笔画落点时间不可靠问题,引入动作单元,通过轻量级门控适配器构建主干并利用无训练分类器梯度消除误差,在StrokeBench上测量,提高了正确放置单个笔画比率,核心帧成为可控轴。
中文摘要 AI 辅助
文本到运动模型能识别提示中的动作,但在笔画落点时间上不可靠,如左右交替的四次拳击很少返回四个可分离的笔画。我们引入名为动作单元(AU)的类型化时间事件,使单个笔画(其身体轨迹、动作类别、时间窗口和冲击时间)成为明确的条件信号。通过轻量级门控适配器在AU集上构建冻结的文本到运动主干,注入两个流(逐笔画令牌和每帧相位通道),并在推理时利用来自冻结帧级检测器的无训练分类器梯度消除残余时间误差。我们在StrokeBench上测量逐笔画控制,其提示指定计数、顺序、轨迹和核心帧位置,并与经过审核的笔画语料库配对。与最强的先前接口相比,AU基础显著提高了正确放置单个笔画的比率,在文本、间隔和帧级基线中具有最佳运动质量。提示的核心帧成为进一步的可控轴。
英文摘要
Text-to-motion models are competent at the action a prompt names but unreliable at when each stroke lands: four punches alternating left and right rarely return four separable strokes. We introduce typed temporal events called Action Units (AUs) that make the individual stroke -- its body track, action class, time window, and impact timing -- an explicit conditioning signal. We ground a frozen text-to-motion backbone on the AU set through a lightweight gated adapter injecting two streams (per-stroke tokens and a per-frame phase channel), and at inference close residual timing errors with a training-free classifier gradient from a frozen frame-level detector. We measure per-stroke control on StrokeBench, whose prompts specify count, ordering, track, and core-frame placement, paired with an audited stroke corpus. AU grounding markedly raises the rate of correctly placed single strokes over the strongest prior interface, at the best motion quality among text-, interval-, and frame-level baselines. The prompted core frame emerges as a further steerable axis.
发表机构
- Seoul National University(首尔国立大学)
机构由 AI 辅助整理,请以论文原文为准。