arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

放大你所注视的:文本到图像生成中的目标显著性增强

Amplify What You Gaze At: Target Saliency Boosting in Text-to-Image Generation

Shengqi Dang, Zhengxi Yu, Feilin Han, Xingyu Lan, Nan Cao

arXiv 2609.30733首次发表:更新:

发表机构

Tongji University; Shanghai Innovation Institute; Fudan University(同济大学; 上海创新研究院; 复旦大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出目标显著性增强任务及 GazeME 框架,通过可学习标记令牌和显著性感知激活策略,在无需视觉先验的情况下提升文本到图像生成中目标对象的视觉显著性,同时保持语义对齐和图像质量。

AI 中文摘要

文本到图像生成在控制对象出现的内容、位置和方式方面已取得进展,但对象之间的视觉注意力分布仍 largely 未被探索。在本文中,我们引入了目标显著性增强(Target Saliency Boosting),这是一项新任务,旨在在文本到图像生成过程中增强特定对象的视觉显著性,而无需任何视觉先验。我们的关键见解是,视觉显著性本质上是相对的:增强目标对象的显著性也依赖于场景中所有对象的全局显著性分布。基于这一见解,我们提出了 GazeME,一个轻量级框架,使用显著性标记提示,在对象描述周围插入可学习的标记令牌,以指示要视觉强调或抑制的对象。为了学习这些标记,我们构建了一个显著性-语义数据集,将图像-提示对中的对象与对象级显著性分数关联起来,并提出了显著性先验标记激活(SPMA),一种显著性感知的随机标记激活策略,利用相对显著性关系进行稳健训练。在推理过程中,GazeME 自动在提示中插入适当的标记,从而直接增强目标对象的视觉显著性。大量实验表明,GazeME 在保持语义对齐和图像质量的同时,有效增强了目标显著性。

英文摘要

Text-to-image generation has advanced in controlling what, where, and how objects appear, yet how visual attention is distributed among objects remains largely unexplored. In this paper, we introduce Target Saliency Boosting, a new task aimed at boosting the visual saliency of a specific object during text-to-image generation without requiring any visual priors. Our key insight is that visual saliency is inherently relative: boosting the saliency of a target object also depends on the global saliency distribution across all objects in the scene. Based on this insight, we propose GazeME, a lightweight framework that uses saliency-marked prompts, inserting learnable marker tokens around object descriptions to indicate which objects to visually emphasize or suppress. To learn these markers, we construct a saliency-semantics dataset that associates objects in image--prompt pairs with object-level saliency scores, and propose Saliency Prior Marker Activation (SPMA), a saliency-aware stochastic marker activation strategy that exploits relative saliency relationships for robust training. During inference, GazeME automatically inserts appropriate markers into the prompt, thereby directly enhancing the visual saliency of the target object. Extensive experiments demonstrate that GazeME effectively boosts target saliency while preserving both semantic alignment and image quality.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑