arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.29123cs.CV

跳舞的简笔画人物:用于训练视频生成模型的入门级数据集

Dancing Stick Figures: An Introductory Dataset for Training Video Generation Models

  • Sprited(斯普里特公司)
  • Dataset(数据集机构)

机构由 AI 辅助整理,请以论文原文为准。

Jin Hyuk Cho

AI总结:

本文针对视频生成模型训练的反馈循环长、数据难获取、评分粗糙三大痛点,构建了跳舞的简笔画人物数据集,该数据集规模适配单GPU训练,提供低预算训练流程,且含丰富标注以支持精准评估。

AI中文摘要:

从零开始训练视频生成模型难度远超模型设计本身,原因有三:一是反馈循环冗长,训练运行后才显现的故障会让每次尝试的修复都需再运行一次训练;二是数据难以获取,强大模型背后的语料库和制作流程规模庞大、异质性强且常不公开;三是评分方式粗糙,开放式生成没有单一正确输出,仅靠汇总分数无法确定样本是否成功或哪一属性存在缺陷。跳舞的简笔画人物(Dancing Stick Figures)是针对这三个障碍构建的合成视频数据集。为实现迭代速度,其64×64、64帧的参考任务规模适合在单张工作站GPU上进行实用的重复训练;为确保可访问性,发布的是0.79GB的训练层级,包含4020个视频片段——1340个6秒的源动作,每个动作由确定性数据集生成工具从三个相机视角渲染而成,还提供了检查点和Colab工作流,可在16GB Tesla T4上以降低的预算重新运行参考训练流程;为实现精准评分,每一帧都保留了其生成状态(ARDY cskel27关节位置、相机、身体参数和源动作)以及逐像素深度、表面法线和部件标签,这些标注支持用于可见拓扑和部件级运动的数据集特定指标,损坏情况可暴露这些指标的敏感性和盲区。

英文摘要:

Training a video-generation model from scratch is hard for reasons that precede model design. The feedback loop is long: a failure that appears only after a training run can make each attempted fix another run. The data are hard to reach: the corpora and recipes behind strong models are large, heterogeneous, and often unreleased. And scoring is blunt: open-ended generation has no single correct output, and an aggregate score does not by itself establish whether a sample succeeds or which property failed. Dancing Stick Figures is a synthetic video dataset built against these three obstacles. For iteration speed, its 64x64, 64-frame reference task is sized for practical repeated training on a single workstation GPU. For accessibility, the release is a 0.79-GB training tier of 4,020 video clips--1,340 six-second source motions, each rendered from three cameras by a deterministic dataset-generation harness--with checkpoints and a Colab workflow that reruns the reference training pipeline at reduced budget on a 16 GB Tesla T4. For scoring, every frame retains its generating state (ARDY cskel27 joint positions, camera, body parameters, and source motion) and per-pixel depth, surface normals, and part labels. These annotations support dataset-specific metrics for visible topology and part-wise motion; corruptions expose their sensitivities and blind spots.

补充信息

↑