MPDS:一个用于扩散模型图像生成的电影海报数据集
MPDS: A Movie Posters Dataset for Image Generation with Diffusion Model
浏览论文内容
中文总结 AI 辅助
本文提出首个电影海报图像-文本数据集 MPDS,包含 373k+ 图像-文本对和 8k+ 演员图像,并设计多条件扩散框架,结合海报提示、海报说明与演员图像,显著提升个性化电影海报生成效果。
中文摘要 AI 辅助
电影海报对于吸引观众、传达主题以及推动电影行业市场竞争至关重要。传统设计费时费力,而智能生成技术在提升效率和改进设计方面具有优势。尽管图像生成取得了令人振奋的进展,现有模型在生成令人满意的海报结果方面仍常常力不从心。主要问题在于缺乏专门的海报数据集用于针对性模型训练。本文提出了一个电影海报数据集(Movie Posters DataSet,MPDS),专为文本到图像生成模型设计,以革新海报制作。作为专门面向海报的数据集,据我们所知,MPDS 是首个图像-文本对数据集,包含 373k+ 图像-文本对和 8k+ 演员图像(覆盖 4k+ 位演员)。详细的海报描述,如电影标题、类型、演员阵容和剧情简介,均基于公开电影简介进行细致整理和标准化,并命名为 movie-synopsis prompt。为了增强海报描述并减少与电影简介之间的差异,我们进一步利用大规模视觉-语言模型自动为每张海报生成视觉感知提示,然后进行人工校正并与 movie-synopsis prompt 整合。此外,我们引入了 poster captions 提示,以展示海报中的文本元素,如演员姓名和电影标题。针对电影海报生成,我们开发了一个多条件扩散框架,以海报提示、海报说明和演员图像(用于个性化)作为输入,通过学习扩散模型获得优异结果。实验表明,我们提出的 MPDS 数据集在推进个性化电影海报生成方面具有重要价值。MPDS 可在 https://anonymous.4open.science/r/MPDS-373k-BD3B 获取。
英文摘要
Movie posters are vital for captivating audiences, conveying themes, and driving market competition in the film industry. While traditional designs are laborious, intelligent generation technology offers efficiency gains and design enhancements. Despite exciting progress in image generation, current models often fall short in producing satisfactory poster results. The primary issue lies in the absence of specialized poster datasets for targeted model training. In this work, we propose a Movie Posters DataSet (MPDS), tailored for text-to-image generation models to revolutionize poster production. As dedicated to posters, MPDS stands out as the first image-text pair dataset to our knowledge, composing of 373k+ image-text pairs and 8k+ actor images (covering 4k+ actors). Detailed poster descriptions, such as movie titles, genres, casts, and synopses, are meticulously organized and standardized based on public movie synopsis, also named movie-synopsis prompt. To bolster poster descriptions as well as reduce differences from movie synopsis, further, we leverage a large-scale vision-language model to automatically produce vision-perceptive prompts for each poster, then perform manual rectification and integration with movie-synopsis prompt. In addition, we introduce a prompt of poster captions to exhibit text elements in posters like actor names and movie titles. For movie poster generation, we develop a multi-condition diffusion framework that takes poster prompt, poster caption, and actor image (for personalization) as inputs, yielding excellent results through the learning of a diffusion model. Experiments demonstrate the valuable role of our proposed MPDS dataset in advancing personalized movie poster generation. MPDS is available at https://anonymous.4open.science/r/MPDS-373k-BD3B.
发表机构
- Nanjing University of Science and Technology(南京理工大学)
- SeetaCloud(视达云)
机构由 AI 辅助整理,请以论文原文为准。