发表机构
Hunyuan, Tencent; Beijing Film Academy; Peking University; Shenzhen University(腾讯混元; 北京电影学院; 北京大学; 深圳大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出首个评估短剧全制作链条各阶段的基准测试集DramaChain Bench,构建配套系统实现模型公平比较与自动评分,证实上游缺陷会级联影响最终剧集质量,为短剧生成研究提供端到端评估支撑。
AI 中文摘要
商业短剧制作遵循多阶段链条:剧本、分镜脚本、关键帧图像、镜头级视频,以及最终完成的短剧。大多数现有基准测试集仅使用预先创作的输入评估视频生成阶段,而非真实的上游流水线输出,这使得两个关键问题无法得到解答:一是各阶段是否符合原始剧本意图(而非仅符合其直接输入提示),二是不同镜头组装成多集内容后是否仍保持连贯性。我们提出 DramaChain Bench,这是首个评估完整制作链条所有阶段的短剧基准测试集。它基于三个共享同一维度体系的内部系统构建:DramaChain Dimensions,即每个阶段都实例化的五个评估轴,分解为63个叶维度。DramaChain Agent 在工作流程和最终短剧质量两方面均对标商业短剧平台,可实现不同模型间的阶段公平比较。DramaChain Labeling System 对5785个条目各由三名专业标注员独立评分,所有缺陷均按空间和时间定位,并从预定义缺陷列表中选取,此过程产生17488个有效评分和255925条可追溯归因记录。人工标注证实上游缺陷会沿流水线级联,表明最终剧集质量并非仅由视频生成决定。DramaChain Agentic Judge 随后自动对每个叶维度评分,在多个智能体轮次中收集证据后,按每个条目的清单进行判断,其以0.918的平均PLCC重现了模型排名,足以无需标注成本即可接纳新模型。
英文摘要
Commercial short-drama production follows a multi-stage chain: script, storyboard, keyframe imagery, shot-level video, and the finished short drama. Most existing benchmarks evaluate solely the video-generation stage using pre-authored inputs instead of real upstream pipeline outputs. This leaves two critical questions unanswerable: whether each stage adheres to the original script intent (rather than only its immediate input prompt), and whether disparate shots remain coherent after assembly into multi-episode releases. We present DramaChain Bench, the first short-drama benchmark that evaluates every stage of the complete production chain. It is built upon three in-house systems sharing one dimension system, DramaChain Dimensions: five evaluation axes instantiated at every stage, resolving into 63 leaf dimensions. DramaChain Agent is calibrated against commercial short-drama platforms in both workflow and finished short-drama quality, enabling stage-wise fair comparison across models. DramaChain Labeling System has each of the 5,785 items scored independently by three professional annotators, with all defects spatio-temporally localised and selected from a predefined defect list. This process produces 17,488 valid scores and 255,925 traceable attribution records. The human annotations confirm that upstream defects cascade across the pipeline, demonstrating that final episode quality is not governed by video generation alone. DramaChain Agentic Judge then scores every leaf dimension automatically, gathering evidence over multiple agentic rounds before judging against a per-item checklist; it reproduces the model ranking at a mean PLCC of 0.918, enough to admit new models at no annotation cost.
Comments50 pages, 19 figures, 19 tables. Technical report. Haoyuan Shi and Mingtao Chen contributed equally. Project lead: Zhichao Hu. Corresponding author: Richeng Xuan