arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

文本驱动的艺术舞台布置:来自绘画的姿态、光照与相机参考

Text-Driven Artistic Staging: 3D Posing, Lighting, and Camera References from Paintings

Yunge Wen

arXiv 2608.28823首次发表:更新:

发表机构

Massachusetts Institute of Technology(麻省理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对现有生成方法独立建模舞台布置元素的问题,提出文本驱动的可编辑3D舞台布置任务,构建相关数据集并训练流匹配Transformer模型,在情感描述的舞台布置检索任务上性能优于CLIP方法,验证了生成3D舞台布置参考的可行性。

AI 中文摘要

艺术家协调人体姿态、光照和相机位置来传递叙事与情感,但现有生成方法通常独立建模这些元素。我们提出文本驱动的可编辑3D舞台布置任务,即从情感描述中联合生成人体姿态、主光源和相机配置。我们通过重构SMPL人体、估计低频光照、恢复相机参数,将2328幅具象绘画构建为11911组文本-舞台布置对,每一场景配对ArtEmis描述。我们训练流匹配Transformer,支持可变数量的人物,为每个提示生成多个舞台布置备选方案。在保留的描述上,该模型达到32.2%的检索R@1,而基于CLIP的最近邻检索为16.6%,同时大致保持语料库级多样性。这些结果证明了从文本生成可编辑、情感条件化的3D舞台布置参考的可行性。

英文摘要

Artists coordinate human pose, illumination, and camera placement to convey narrative and emotion, but existing generative methods typically model these elements independently. We introduce text-to-editable 3D staging, a task that jointly generates human poses, a dominant light, and a camera configuration from an affective description. We construct 11,911 text--staging pairs from 2,328 figurative paintings by reconstructing SMPL bodies, estimating low-frequency illumination, recovering camera parameters, and pairing each scene with ArtEmis descriptions. We train a flow-matching transformer that supports variable numbers of figures and produces multiple staging alternatives for each prompt. On held-out descriptions, the model achieves 32.2\% retrieval R@1, compared with 16.6\% for CLIP-based nearest-neighbor retrieval, while approximately preserving corpus-level diversity. These results demonstrate the feasibility of generating editable, emotionally conditioned 3D staging references from text.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑