超越《星夜》:面向艺术家基础的文本到图像生成的捷径感知控制状态规划
Beyond Starry Night: Shortcut-Aware Control-State Planning for Artist-Grounded Text to Image Generation
浏览论文内容
中文总结 AI 辅助
该研究针对艺术家基础的文本到图像生成存在的捷径偏差问题,提出 Atelier 框架,结合 ArtIntentBench 基准测试,提升了风格保真度与结构保留度,减少了捷径替换。
中文摘要 AI 辅助
面向艺术家基础的图像生成,仅在提示词中添加艺术家名称是不够的。图像模型常通过规范捷径响应艺术家名称,如重复出现的 motifs(主题元素)、通用调色板或过度代表的时期特征,而非保留用户预期的场景。我们推出 Atelier,这是一个面向艺术家基础的图像生成的捷径感知控制状态规划框架。Atelier 将未明确的艺术意图转化为显式控制状态,该状态分离场景锚点、保留/转换决策、风格流派假设、角色绑定的艺术家证据以及捷径规避约束。它利用艺术家层面知识和局部补丁参考来锚定此状态,编译后端感知的生成计划,并通过全局和局部真实性反馈迭代优化候选。我们还推出 ArtIntentBench,这是一个涵盖梵高和齐白石的基准,涉及艺术品重新渲染、时期/风格控制生成、历史上未出现的主题、捷径审计以及人类偏好评估。在开源和闭源生成器中,与提示工程、检索增强和通用智能体基线相比,Atelier 提高了艺术家层面的风格保真度,更忠实地保留了源结构,并大幅减少了捷径替换。这些结果表明,艺术家基础的生成不仅受图像合成的瓶颈限制,还受上游显式、证据基础的艺术控制的推理瓶颈限制。
英文摘要
Artist-grounded image generation requires more than appending an artist name to a prompt. Image models often respond to artist names through canonical shortcuts, such as recurring motifs, generic palettes, or overrepresented period signatures, rather than preserving the user's intended scene. We introduce Atelier, a shortcut-aware control-state planning framework for artist-grounded image generation. Atelier translates underspecified artistic intent into an explicit control state that separates scene anchors, preserve/transform decisions, style-regime hypotheses, role-bound artist evidence, and shortcut-avoidance constraints. It grounds this state using artist-level knowledge and local patch references, compiles backend-aware generation plans, and iteratively refines candidates through global and local authenticity feedback. We further introduce ArtIntentBench, a benchmark covering Van Gogh and Qi Baishi across artwork re-rendering, period/style-controlled generation, historically unseen subjects, shortcut auditing, and human preference evaluation. Across open-weight and closed-source generators, Atelier improves artist-level style fidelity, preserves source structure more faithfully, and substantially reduces shortcut substitution compared with prompt-engineered, retrieval-augmented, and general-purpose agent baselines. These results suggest that artist-grounded generation is bottlenecked not only by image synthesis, but by the upstream inference of explicit, evidence-grounded artistic controls.
发表机构
- Jilin University(吉林大学)
- Adobe(奥多比公司)
- University of Wisconsin(威斯康星大学)
机构由 AI 辅助整理,请以论文原文为准。