arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.06751cs.CVcs.AI

超越《星夜》:面向艺术家基础的文本到图像生成的捷径感知控制状态规划

Beyond Starry Night: Shortcut-Aware Control-State Planning for Artist-Grounded Text to Image Generation

Kuan Xing, Ye Wang, Changyi Gan, Yuheng Li, Thao Nguyen, Yi Chang, Yilin Wang

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对艺术家基础的文本到图像生成存在的捷径偏差问题,提出 Atelier 框架,结合 ArtIntentBench 基准测试,提升了风格保真度与结构保留度,减少了捷径替换。

中文摘要 AI 辅助

面向艺术家基础的图像生成,仅在提示词中添加艺术家名称是不够的。图像模型常通过规范捷径响应艺术家名称,如重复出现的 motifs(主题元素)、通用调色板或过度代表的时期特征,而非保留用户预期的场景。我们推出 Atelier,这是一个面向艺术家基础的图像生成的捷径感知控制状态规划框架。Atelier 将未明确的艺术意图转化为显式控制状态,该状态分离场景锚点、保留/转换决策、风格流派假设、角色绑定的艺术家证据以及捷径规避约束。它利用艺术家层面知识和局部补丁参考来锚定此状态,编译后端感知的生成计划,并通过全局和局部真实性反馈迭代优化候选。我们还推出 ArtIntentBench,这是一个涵盖梵高和齐白石的基准,涉及艺术品重新渲染、时期/风格控制生成、历史上未出现的主题、捷径审计以及人类偏好评估。在开源和闭源生成器中,与提示工程、检索增强和通用智能体基线相比,Atelier 提高了艺术家层面的风格保真度,更忠实地保留了源结构,并大幅减少了捷径替换。这些结果表明,艺术家基础的生成不仅受图像合成的瓶颈限制,还受上游显式、证据基础的艺术控制的推理瓶颈限制。

英文摘要

Artist-grounded image generation requires more than appending an artist name to a prompt. Image models often respond to artist names through canonical shortcuts, such as recurring motifs, generic palettes, or overrepresented period signatures, rather than preserving the user's intended scene. We introduce Atelier, a shortcut-aware control-state planning framework for artist-grounded image generation. Atelier translates underspecified artistic intent into an explicit control state that separates scene anchors, preserve/transform decisions, style-regime hypotheses, role-bound artist evidence, and shortcut-avoidance constraints. It grounds this state using artist-level knowledge and local patch references, compiles backend-aware generation plans, and iteratively refines candidates through global and local authenticity feedback. We further introduce ArtIntentBench, a benchmark covering Van Gogh and Qi Baishi across artwork re-rendering, period/style-controlled generation, historically unseen subjects, shortcut auditing, and human preference evaluation. Across open-weight and closed-source generators, Atelier improves artist-level style fidelity, preserves source structure more faithfully, and substantially reduces shortcut substitution compared with prompt-engineered, retrieval-augmented, and general-purpose agent baselines. These results suggest that artist-grounded generation is bottlenecked not only by image synthesis, but by the upstream inference of explicit, evidence-grounded artistic controls.

发表机构

  • Jilin University(吉林大学)
  • Adobe(奥多比公司)
  • University of Wisconsin(威斯康星大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑