arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

T2LSC-Bench:文本到图像生成中的本地化语义控制基准

T2LSC-Bench: Benchmarking Localized Semantic Control in Text-to-Image Generation

Yan Wang, Xinyi Hou, Weiguo Lin, Junjun Si, Siwei Ma

arXiv 2609.02255首次发表:更新:

发表机构

Communication University of China; Huazhong University of Science and Technology; Peking University(中国传媒大学; 华中科技大学; 北京大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出用于评估文本到图像生成本地化语义控制的基准T2LSC-Bench,发现现有模型存在语义泄漏问题,反泄漏提示可降低泄漏率且不影响渲染准确率。

AI 中文摘要

近期文本到图像模型渲染明确文本的能力不断提升,但可靠的本地化文本控制仅生成正确字符串是不够的。在产品标签、标识、界面设计等应用场景中,目标文本需渲染在指定的文本承载区域内,且不得改变预设的主体身份或周围场景语义。我们将违反该要求的情况称为目标文本关联的语义泄漏,即目标文本的语义通过指定锚点之外的非文本视觉内容体现。现有视觉-文本基准主要评估可读性、拼写准确性和布局,基本未涉及这类语义泄漏问题。我们提出T2LSC-Bench,这是一个可控诊断基准,包含50个种子主体,每个模型对应1200个提示案例,共评估6个模型的7160张图像。其因子化设计会改变语义关系、场景开放性、提示模式和语言。双分支协议结合OCR-VLM文本验证与结构化VLM语义判断,用于测量锚点文本准确率(TAA)、语义主体保留率(SSP)、语义泄漏率(SLR)和条件语义泄漏率(cSLR)。在压力测试条件下,SLR从1.2%升至18.1%,cSLR从1.3%升至18.2%,而TAA仅从91.4%降至90.9%;反泄漏提示可将SLR从16.6%降至8.4%,且不降低渲染准确率。对420张图像的人工验证显示,自动标注与人工裁决标注的一致性较强。这些结果表明,准确的文本渲染无法保证目标文本语义的本地化包含。

英文摘要

Recent text-to-image models have become increasingly capable of rendering explicit text, but reliable localized text control requires more than generating the correct string. In applications such as product labeling, signage, and interface design, target text should be rendered within a designated text-bearing region without altering the predefined subject identity or surrounding scene semantics. We refer to violations of this requirement as target-text-associated semantic leakage, in which target-text semantics are expressed through non-textual visual content beyond the designated anchor. Existing visual-text benchmarks primarily evaluate readability, spelling accuracy, and layout, leaving this form of semantic leakage largely unexamined. We introduce T2LSC-Bench, a controlled diagnostic benchmark comprising 50 seed subjects and 1,200 prompt cases per model, yielding 7,160 evaluated images across six models. Its factorized design varies semantic relation, scene openness, prompt mode, and language. A dual-branch protocol combines OCR-VLM text verification with structured VLM semantic judgments to measure Text-at-Anchor Accuracy (TAA), Semantic Subject Preservation (SSP), Semantic Leakage Rate (SLR), and Conditional Semantic Leakage Rate (cSLR). Under stress-test conditions, SLR increases from 1.2% to 18.1% and cSLR from 1.3% to 18.2%, whereas TAA decreases only from 91.4% to 90.9%. Anti-leakage prompting reduces SLR from 16.6% to 8.4% without degrading rendering accuracy. Human validation on 420 images shows strong agreement between automatic and adjudicated annotations. These results show that accurate text rendering does not guarantee local containment of target-text semantics.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑