发表机构
Nanyang Technological University(南洋理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对固定时间网格难以表征运动多尺度动态的问题,提出FreqMo,通过小波频带分解与统一频率残差量化实现频率解耦,压缩三倍并提升高频保真度,达到SOTA效果。
AI 中文摘要
大多数人体运动生成方法将运动编码为均匀时间网格上的令牌,其中每个令牌覆盖相同的固定时间窗口。然而,人体运动在时间上是异质的:缓慢演化的全局轨迹与快速瞬态事件(如足部接触和关节冲量)共存。将这种多尺度动态强制施加到具有相同时间分辨率的令牌上,会纠缠运动频率,导致缓慢区域冗余,同时平滑掉区分真实运动的关键快速细节。我们提出FreqMo,一种尺度自适应的运动表示方法,它将运动分解为小波频带,在时间尺度上分离动态,同时保持时间定位和精确重建。统一频率残差量化(UFRQ)随后将所有频带编码到单个共享码本中,将令牌序列压缩三倍,并实现稳定的单阶段生成。实验表明,FreqMo达到了最先进的保真度,并显著改善了高频保留,且相同的分解可迁移到连续扩散骨干网络。
英文摘要
Most human motion generation methods encode motion as tokens on a uniform temporal grid, where every token spans the same fixed time window. Human motion, however, is temporally heterogeneous: slowly evolving global trajectories coexist with rapid transient events such as foot contacts and joint impulses. Forcing such multi-scale dynamics onto tokens of identical temporal resolution entangles motion frequencies, leaving slow regions redundant while smoothing out the rapid details that distinguish realistic motion. We propose \textbf{FreqMo}, a scale-adaptive motion representation that decomposes motion into wavelet frequency bands, separating dynamics across temporal scales while preserving temporal localization and exact reconstruction. Unified Frequency Residual Quantization (UFRQ) then encodes all bands within a single shared codebook, compressing the token sequence threefold and enabling stable single-stage generation. Experiments show FreqMo attains SOTA fidelity with substantially improved high-frequency preservation, and the same decomposition transfers to continuous diffusion backbones.