发表机构
School of Artificial Intelligence, The Chinese University of Hong Kong, Shenzhen; vivo AI Lab; University of the Chinese Academy of Sciences; Southeast University; Peking University; School of Science and Engineering, The Chinese University of Hong Kong, Shenzhen; Shenzhen Future Network of Intelligence Institute (FNii-Shenzhen); Guangdong Provincial Key Laboratory of Future Networks of Intelligence, CUHK-Shenzhen(香港中文大学(深圳)人工智能学院; vivo AI实验室; 中国科学院大学; 东南大学; 北京大学; 香港中文大学(深圳)理工学院; 深圳未来智能网络研究院; 香港中文大学(深圳)广东省未来智能网络重点实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对可执行视频编辑规划问题,提出基于验证器的RefineCut及RefineCut-Evo方法,训练8B开放权重规划器,在基准上取得优于前沿模型的效果,代码与基准已公开。
AI 中文摘要
实用视频编辑不仅是像素生成:编辑者必须将脚本、素材池、音乐元数据和硬约束转化为可执行时间线。我们将该决策层研究为「可执行视频编辑规划」,并推出RefineCut,与包装提示前沿模型的工作流系统不同,它为此训练了一个紧凑的开放权重规划器。该规划器通过结构化补丁编辑类型化时间线,涵盖素材选择、修剪、排序、转场、时长及音乐对齐;确定性验证器应用每个补丁并对照显式约束账本检查。由于编辑无单一真实修复,我们不直接模仿教师:RefineCut通过验证器重放所有多教师分支,保留验证器最优修复作为监督。第二阶段RefineCut-Evo让学生用验证器和任务准则评分自身修复,并基于高边际偏好对训练,最终8B规划器在推理时以封闭验证器循环运行,无需教师调用。在RefineCut-Bench(3578个任务、7971个带字幕素材、499条音乐曲目、显式账本)上,验证器重放蒸馏使规划器在协议特定视频编辑得分从0.620提升至0.858,RefineCut-Evo达到0.924;该增益可迁移至Llama-3.1-8B和GLM-4-9B,在同一封闭循环中,8B规划器表现匹配或超越其前沿教师。代码和RefineCut-Bench已公开,详见数据可用性声明。
英文摘要
Practical video editing is not only pixel generation: an editor must turn a brief, a clip pool, music metadata, and hard constraints into an executable timeline. We study this decision layer as \emph{executable video-editing planning} and introduce RefineCut, which, unlike workflow systems that wrap a prompted frontier model, trains a compact open-weight planner for it. The planner edits a typed timeline through structured patches covering clip selection, trimming, ordering, transitions, and duration and music alignment; a deterministic verifier applies each patch and checks it against an explicit constraint ledger. Because editing has no single ground-truth repair, we do not imitate teachers directly: RefineCut replays every multi-teacher branch through the verifier and keeps verifier-best repairs as supervision. A second stage, RefineCut-Evo, lets the student score its own repairs with the verifier and a task rubric and trains on high-margin preference pairs, so the final $8$B planner runs in a closed verifier loop with no teacher calls at inference. On RefineCut-Bench ($3{,}578$ tasks, $7{,}971$ captioned clips, $499$ music tracks, explicit ledgers), verifier-replayed distillation lifts the planner from $0.620$ to $0.858$ on the protocol-specific Video-Editing Score and RefineCut-Evo reaches $0.924$; the gain transfers to Llama-3.1-8B and GLM-4-9B, and in the same closed loop the $8$B planner matches or exceeds its frontier teachers. Code and RefineCut-Bench are publicly released; see the Data Availability statement.
CommentsAccepted to the Main Conference of EMNLP '26