arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从视频片段到创作轨迹:面向AI原生视频创作的Sora100K

From Video Clips to Creation Trajectory: Sora100K for AI-Native Video Creation

Sicong Yang, Ruihuan Yang, Jian Lu, Jianfei Yuan, Xiaodong Cun, Xiuli Bi

arXiv 2610.11770首次发表:更新:

发表机构

Chongqing University of Posts and Telecommunications; Great Bay University(重庆邮电大学; 大湾区大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出面向AI原生视频创作的Sora100K数据集,将创作工作流表示为结构化轨迹,经标注与质量控制后适配LTX-2模型,可提升视频生成相关性能,为该领域提供新数据基础。

AI 中文摘要

AI原生视频创作正从孤立的视频片段向迭代式视频创作工作流转变。然而,现有数据集大多仍为视频片段,将视频生成与编辑视为独立任务,而非视频创作工作流的关联阶段。本文提出Sora100K,该数据集将AI原生视频创作工作流表示为结构化视频创作轨迹。具体而言,我们首先识别视频创作轨迹,并根据其结构角色将其分解为三个子集:文本到视频生成记录作为根节点、单轮视频编辑记录作为编辑边、多轮视频编辑记录作为完整轨迹。随后,我们使用VLM为生成根节点分配语义标注,为编辑边分配编辑操作标注。严格的构建流程进一步重构了源到编辑的谱系、编辑顺序和中间视频状态,同时确保数据质量。最后,我们对LTX-2模型进行轻量适配,以评估Sora100K的监督价值。结果显示,该数据集在视觉质量、多镜头生成和跨镜头一致性方面均有提升,而多轮评估表明,遵循多轮编辑指令仍具挑战性。Sora100K为AI原生视频创作建立了超越孤立视频片段、迈向结构化视频创作轨迹的新数据基础。该数据集及补充材料可在指定URL公开获取。

英文摘要

AI-Native video creation is shifting from isolated video clips toward iterative video creation workflows. However, existing datasets remain largely video clips, representing video generation and editing as separate tasks rather than connected stages of a video creation workflow. In this paper, we introduce Sora100K, a dataset that represents the AI-Native video creation workflow as a structured video creation trajectory. Specifically, we first identify video creation trajectories and decompose them into three subsets according to their structural roles: text-to-video generation records as roots, single-turn video editing records as editing edges, and multi-turn video editing records as complete trajectories. Then, we use a VLM to assign semantic annotations for generation roots and editing-operation annotations for editing edges. A strict construction pipeline further reconstructs source-to-edit lineage, editing order, and intermediate video states while ensuring data quality. Finally, we perform lightweight adaptation on LTX-2 models to assess the supervision value of Sora100K. The results show improvements in visual quality, multi-shot generation, and cross-shot consistency, while successive-turn evaluation reveals that following multi-turn editing instructions remains challenging. Sora100K establishes a new data foundation for AI-Native video creation beyond isolated video clips and toward structured video creation trajectory. The dataset and supplementary materials are publicly available at https://huggingface.co/datasets/ysicong/Sora100K.

Comments18 pages, 17 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑