Poplar:面向以人为中心的图像数据集合成的可扩展流水线
Poplar: A Scalable Pipeline for Human-Centric Image Dataset Synthesis
浏览论文内容
中文总结 AI 辅助
本文提出可扩展流水线Poplar,通过“指定-渲染-检查”流程构建高质量以人为中心图像数据集,产出含9401对样本的Poplar-9K数据集并开源相关资源。
中文摘要 AI 辅助
现有图像生成器可合成逼真的以人为中心的图像,但生成有用的数据集与生成单张成功图像不同。以人为中心的数据集必须涵盖多样的人物与场景,避免不合理的属性组合,保留日常摄影特征,并规模化暴露质量控制决策。我们提出Poplar,一种可复现的“指定-渲染-检查”以人为中心图像数据集合成流水线:指定阶段在常识约束下采样结构化属性,将其表述为面向摄影的提示;渲染阶段采用适配逼真度的图像生成器,结合感知构图的宽高比设置,对明显技术故障进行重试;检查阶段对每个候选样本应用单一结构化视觉-语言审核,保留原始提示,同时拒绝图像固有缺陷或与提示不匹配的样本。借助Poplar,我们构建了Poplar-9K:从11765个已审核候选样本中保留的9401个精选以人为中心的图像-文本对,接受率为79.9%。我们将该数据集连同流水线、配置、不可变生成提示及可审计检查记录一同发布,作为构建可定制以人为中心数据集的紧凑资源。
英文摘要
Recent image generators can synthesize convincing human-centric images, yet producing a useful collection remains different from producing a single successful image. A human-centric dataset must cover varied people and contexts, avoid implausible attribute combinations, preserve an everyday photographic character, and expose quality-control decisions at scale. We present Poplar, a reproducible Specify--Render--Inspect pipeline for human-centric image dataset synthesis. Specify samples structured attributes under commonsense constraints and verbalizes them as photography-oriented prompts. Render uses a realism-adapted image generator across composition-aware aspect ratios and retries obvious technical failures. Inspect applies a single structured vision--language review to each candidate, preserving the original prompt while rejecting intrinsic image defects or material prompt mismatches. Using Poplar, we construct Poplar-9K: 9,401 curated human-centric image--text pairs retained from 11,765 reviewed candidates (79.9\% acceptance). We release the dataset together with the pipeline, configurations, immutable generation prompts, and auditable inspection records as a compact resource for building customizable human-centric collections.
发表机构
- Beijing University of Posts and Telecommunications(北京邮电大学)
机构由 AI 辅助整理,请以论文原文为准。