发表机构
Hokkaido University(北海道大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对人体运动预测中数据集蒸馏因缺乏先验导致合成运动不真实的问题,提出潜在蒸馏框架,利用RVQ-VAE压缩运动并冻结解码器约束输出,在多个数据集上优于基线。
AI 中文摘要
数据集蒸馏(DD)将大规模训练集压缩为紧凑的合成集,同时保留下游训练效用。尽管DD已在图像领域得到广泛研究,并最近扩展到时间序列预测,但其在人体运动预测中的应用仍基本未被探索。人体运动具有高维性和结构耦合性,原始运动空间中的梯度匹配(GM)优化了许多相关变量,却缺乏对姿态合理性或时间动态的先验知识,这常常导致合成运动不切实际且不稳定。为解决这一局限,我们提出了一种潜在DD框架,利用学习到的运动先验对蒸馏进行正则化。运动首先由残差量化变分自编码器(RVQ-VAE)压缩,然后蒸馏仅通过冻结的量化器和解码器更新可学习的潜在库。预训练解码器将合成运动限制在其输出空间内,而残差量化通过多个码本逐步细化潜在近似,缓解了单阶段向量量化的表示瓶颈。在Human3.6M、CMU和3DPW数据集上使用两种预测主干的实验表明,所提出的框架在30个评估设置中的27个中优于直接GM,并在所有设置中优于随机子集,且在定性比较中产生明显更合理的合成运动。
英文摘要
Dataset distillation (DD) compresses a large training set into a compact synthetic set while preserving downstream training utility. Although DD has been widely studied for images and recently extended to time-series forecasting, its application to human motion prediction remains largely unexplored. Human motion is high-dimensional and structurally coupled, and gradient matching (GM) in the original motion space optimizes many correlated variables without a prior on pose plausibility or temporal dynamics, which frequently yields implausible and unstable synthetic motions. To address this limitation, we propose a latent DD framework that regularizes distillation with a learned motion prior. Motions are first compressed by a residual-quantized variational autoencoder (RVQ-VAE), and distillation then updates only a learnable latent bank through the frozen quantizer and decoder. The pretrained decoder restricts synthetic motions to its output space, while residual quantization progressively refines the latent approximation across multiple codebooks and alleviates the representational bottleneck of single-stage vector quantization. Experiments on Human3.6M, CMU, and 3DPW with two prediction backbones show that the proposed framework outperforms direct GM in 27 of 30 evaluated settings and random subsets in every setting, and produces visibly more plausible synthetic motions in qualitative comparisons.