发表机构
Kuaishou(快手)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对短视频一键创作需求,提出多智能体系统TemplateCraft,通过规划、素材生成与效果工作流编译将自然语言转为可执行模板,在TemplateBench上显著提升生成成功率与风格一致性。
AI 中文摘要
短视频的日益流行推动了对一键式内容创作的需求。视觉模板通过预设效果将上传的图像转化为个性化内容,但可复用模板的生成在资产准备和工具编排方面仍需要大量人工操作。我们提出了TemplateCraft,一个多智能体系统,通过规划、素材生成、效果工作流生成和协议编译,将自然语言指令转化为客户端可执行的模板。其规划器-评估器循环利用执行反馈进行定向回滚,而阶段级和长期记忆支持在不更新参数的情况下进行修订。我们在从60个真实模板中衍生出的TemplateBench上评估了TemplateCraft。在相同的Qwen3-VL主干下,与仅规划器(三次取最优)相比,TemplateCraft将图像/视频生成成功率从56.7%/30.0%提升至66.7%/50.0%,并提高了模板遵循度和风格一致性。通过额外的评估和修订,它在选定指标上达到或超过了仅规划器的GPT-4o基线。持久化资产进一步提高了跨输入的风格一致性。
英文摘要
The growing popularity of short videos has driven demand for one-click content creation. Visual templates turn uploaded images into personalized content with preset effects, but reusable template generation still requires substantial manual effort in asset preparation and tool orchestration. We propose TemplateCraft, a multi-agent system that converts natural-language instructions into client-executable templates through planning, material generation, effect-workflow generation, and protocol compilation. Its Planner-Evaluator loop uses execution feedback for targeted rollback, while stage-level and long-term memory support revision without parameter updates. We evaluate TemplateCraft on TemplateBench, derived from 60 real-world templates. With the same Qwen3-VL backbone, TemplateCraft raises image/video generation success rates from 56.7%/30.0% to 66.7%/50.0% over Planner-only (best-of-three) and improves template adherence and style consistency. With additional evaluation and revision, it matches or exceeds a GPT-4o Planner-only baseline on selected metrics. Persistent assets further improve cross-input style consistency.
Comments5 pages, 3 figures, 1 table. Submitted to ICASSP 2027