发表机构
Massachusetts Institute of Technology(麻省理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对现有生成方法独立建模舞台布置元素的问题,提出文本驱动的可编辑3D舞台布置任务,构建相关数据集并训练流匹配Transformer模型,在情感描述的舞台布置检索任务上性能优于CLIP方法,验证了生成3D舞台布置参考的可行性。
AI 中文摘要
艺术家协调人体姿态、光照和相机位置来传递叙事与情感,但现有生成方法通常独立建模这些元素。我们提出文本驱动的可编辑3D舞台布置任务,即从情感描述中联合生成人体姿态、主光源和相机配置。我们通过重构SMPL人体、估计低频光照、恢复相机参数,将2328幅具象绘画构建为11911组文本-舞台布置对,每一场景配对ArtEmis描述。我们训练流匹配Transformer,支持可变数量的人物,为每个提示生成多个舞台布置备选方案。在保留的描述上,该模型达到32.2%的检索R@1,而基于CLIP的最近邻检索为16.6%,同时大致保持语料库级多样性。这些结果证明了从文本生成可编辑、情感条件化的3D舞台布置参考的可行性。
英文摘要
Artists coordinate human pose, illumination, and camera placement to convey narrative and emotion, but existing generative methods typically model these elements independently. We introduce text-to-editable 3D staging, a task that jointly generates human poses, a dominant light, and a camera configuration from an affective description. We construct 11,911 text--staging pairs from 2,328 figurative paintings by reconstructing SMPL bodies, estimating low-frequency illumination, recovering camera parameters, and pairing each scene with ArtEmis descriptions. We train a flow-matching transformer that supports variable numbers of figures and produces multiple staging alternatives for each prompt. On held-out descriptions, the model achieves 32.2\% retrieval R@1, compared with 16.6\% for CLIP-based nearest-neighbor retrieval, while approximately preserving corpus-level diversity. These results demonstrate the feasibility of generating editable, emotionally conditioned 3D staging references from text.