发表机构
Fudan University; Shanghai University of Electric Power(复旦大学; 上海电力大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究揭示图像生成模型中渲染文本的语义泄漏现象,提出缓解方法,发现其为可见数据与潜在语义的两用载体,暴露跨模态攻击面。
AI 中文摘要
图像生成模型(IGM)的可靠性与可问责性是构建负责任、可信赖AI系统的关键。近期的IGM,如Nano Banana和GPT-Image,已支持复杂指令遵循、逼真图像合成及可控场景文本渲染。随着这些能力的扩展,安全分析必须考虑复杂提示引入的新控制通道。本研究聚焦渲染文本语义泄漏,这是开放域文本渲染中一个被严重忽视的现象:尽管渲染文本旨在作为应逐字复现的局部视觉约束,但其携带的语言语义可能被模型解读为输入指令的一部分,使渲染文本成为潜在的语义控制通道,其安全影响尚未得到充分理解。我们通过将主视觉提示与渲染文本解耦,测量二者对生成图像的单独及组合影响,系统表征该现象;量化语义泄漏与渲染保真度,进一步分析泄漏如何从模型中间证据中产生。随后,我们证明嵌入场景文本的有害语义可通过基于LLM的提示增强管道持续存在,并引导非文本图像区域,即使主视觉提示保持无害。最后,我们提出一种初步缓解方法,在保留FLUX-2-dev上预期文本渲染行为的同时,减少从渲染文本到非文本区域的不安全语义传递。研究发现揭示渲染文本是可见数据与潜在语义的两用载体,暴露了现代IGM中以文本为中心的跨模态攻击面。
英文摘要
The reliability and accountability of image generative models (IGMs) are essential for building responsible and trustworthy AI systems. Recent IGMs, such as Nano Banana and GPT-Image, now support complex instruction following, realistic image synthesis, and controllable scene-text rendering. As these capabilities expand, safety analysis must also account for new control channels introduced by complex prompts. In this work, we study rendered-text semantic leakage, a largely overlooked phenomenon in open-domain text rendering. Although rendered text is intended to serve as a local visual constraint that should be reproduced verbatim in the generated image, it also carries linguistic semantics that may be interpreted by the model as part of the input instruction. This makes rendered text a potential semantic control channel whose safety implications remain insufficiently understood. We systematically characterize this phenomenon by decoupling the main visual prompt from the rendered text and measuring their individual and compositional effects on generated images. We quantify semantic leakage and rendering fidelity, and further analyze how leakage emerges from intermediate model evidence. We then show that harmful semantics embedded in scene text can persist through LLM-based prompt enhancement pipelines and steer non-text image regions, even when the main visual prompt remains benign. Finally, we propose a preliminary mitigation approach that reduces unsafe semantic transfer from rendered text to non-text regions while preserving the intended text-rendering behavior on FLUX-2-dev. Our findings reveal rendered text as a dual-use carrier of visible data and latent semantics, exposing a text-centric cross-modal attack surface in modern IGMs.
CommentsTo appear in the 2027 IEEE Symposium on Security and Privacy (IEEE S&P 2027)