arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DOGS:面向提示驱动的标志生成的设计空间采样

DOGS: Design-Space Sampling for Prompt-Driven Logo Generation

Ganyu Zou, Chen Dai, Nathan Self, Kevin Piper, Ramachandra Rao Seethiraju, Karthik Shyamsunder, Chang-Tien Lu, Naren Ramakrishnan

arXiv 2610.10760首次发表:更新:

发表机构

Virginia Tech; Verisign, Inc.(弗吉尼亚理工大学; 威瑞信公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对标志生成任务提出DOGS方法,将标志提示转化为结构化设计空间采样,结合原创性感知GFlowNet采样器,生成的标志兼具高可识别性、美学性与多样性,且能有效规避商标侵权。

AI 中文摘要

针对文本到图像(T2I)生成的提示优化几乎完全以文本重写的形式开展,即把简短的用户提示扩展为更长的、模型偏好的标记序列。我们认为这种语言空间的表述并不适合标志创建这类结构化视觉设计任务,因为一行提示会留下大部分设计决策未被指定。这些决策依赖于线性序列无法编码的关系先验,还会留下一个不受控制的渠道,可能导致受保护标志被复制。因此,我们将标志提示重新定义为在结构化设计空间内的采样,并将该想法实例化为DOGS(具有原创性感知GFlowNet采样器的设计空间提示)。我们从大量真实世界标志语料库中挖掘出一种类型化的、图结构的设计语法,其边记录了经验共现关系。随后,GFlowNet采样器生成设计图,其概率与终端奖励成正比,该奖励结合了可识别性、美学以及相对于语料库的原创性。每个槽位仅从封闭的设计级词汇中选取,且在解析过程中会移除任何可能引发侵权或有害的标记。原创性奖励进一步惩罚与现有标志的接近程度,从而在方法构建中纳入侵权规避。在两个开源渲染器上与九个基线对比,DOGS生成的标志更具可识别性和美学性,多样性显著更高,且更不易出现商标侵权问题。

英文摘要

Prompt optimization for text-to-image (T2I) generation has been pursued almost entirely as text rewriting, in which a short user brief is expanded into a longer, model-preferred token sequence. We argue that such a language-space formulation is ill-suited to structured visual design tasks such as logo creation, where a one-line brief leaves most design decisions unspecified. These decisions depend on relational priors that a linear sequence cannot encode, and they leave an uncontrolled channel through which protected marks may be reproduced. We therefore recast logo prompting as sampling within a structured design space, and instantiate this idea as DOGS (Design-space prompting with an Originality-aware GFlowNet Sampler). From a large corpus of real-world logos, we mine a typed, graph-structured design grammar whose edges record empirical co-occurrence. A GFlowNet sampler then generates design graphs with probability proportional to a terminal reward that combines recognizability, aesthetics, and corpus-relative originality. Every slot draws only from a closed design-level vocabulary, and any infringement-inducing or harmful token is removed during parsing. The originality reward further penalizes proximity to existing logos, thereby incorporating infringement avoidance into the method by construction. On two open-source renderers and against nine baselines, DOGS produces logos that are more recognizable and aesthetic, substantially more diverse, and far less prone to trademark infringement.

CommentsAccepted to BMVC 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑