arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

计划而非解码器:诊断与修复推理增强型文本到图像生成中的组合式失败

The Plan, Not the Decoder: Diagnosing and Repairing Compositional Failure in Reasoning-Augmented Text-to-Image Generation

Ashritha Gonuguntla

arXiv 2608.21713首次发表:更新:

发表机构

Carnegie Mellon University(卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对推理增强型文本到图像生成的组合式失败,发现问题源于规划器而非解码器,通过编辑计划(如重写几何结构)可显著提升生成质量,还发布了相关评估协议及数据。

AI 中文摘要

推理增强型文本到图像模型(如GoT-R1)在生成图像 token 前会输出明确的文本计划——包含对象名称、属性和边界框。当此类模型在组合式提示下失败时,问题出在计划错误,还是计划正确但解码器执行不忠实?由于计划是机器可读的,可在解码前对其进行编辑,这使得两者可分离。我们首先验证该标尺:在模型自身的链中交换两个边界框,生成的布局明显翻转,基于检测器的准确率从0.75降至0.48(p<1e-3),而广泛使用的基于VQA的空间指标上升。一项由5名评分者参与的人类研究,与检测器的一致性达81%,与VQA评判者的一致性达57%。因此,所有空间结果均采用几何评分。在合理测量下,解码器是忠实的执行者:94%的生成布局实现了计划的关系,且对象-框绑定在计划的对象片段重新排序后仍能保留。规划器是瓶颈:它因措辞相关原因写出错误关系——语义相同的布局中,“左侧”的准确率为98%,而“右侧”仅为54%,我们通过提及顺序控制分离出这种光栅顺序偏差,且会生成解码器忠实再现的杂乱几何结构。因此,编辑计划无需重新训练即可修复图像:通过重采样进行符号验证可提升5.0个点(p<1e-3),最小原位修复提升6.0个点(p=0.02),仅重写框几何结构提升10.7个点(p<1e-4),完全替换计划提升13.3个点(p=1e-4)。提升幅度与计划的散文风格及规划器下的可能性无关,但与几何结构相关。因此,模块化规划器-解码器设计是可行的,前提是计划内部一致:框-文本矛盾会导致对象重复和身份融合。我们发布了计划-保真度评估协议、所有计划及12000张生成图像。

英文摘要

Reasoning-augmented text-to-image models such as GoT-R1 emit an explicit textual plan - object names, attributes, and bounding boxes - before generating image tokens. When such a model fails a compositional prompt, is the plan wrong, or is the plan right and the decoder unfaithful? Because the plan is machine-readable it can be edited before decoding, which makes the two separable. We first validate the ruler. Swapping the two bounding boxes inside the model's own chain demonstrably flips the generated layout: detector-based accuracy falls 0.75 -> 0.48 (p<1e-3), while a widely used VQA-based spatial metric rises. A five-rater human study agrees with the detector on 81% of items and with the VQA judge on 57%. All spatial results therefore use geometric scoring. Under sound measurement the decoder is a faithful executor: 94% of generated layouts realize the planned relation, and object-box binding survives reordering of the plan's object segments. The planner is the bottleneck. It writes wrong relations for phrasing-dependent reasons - 98% accuracy on "left" against 54% on "right" for semantically identical layouts, a raster-order bias we isolate with a mention-order control - and cluttered geometry that the decoder faithfully reproduces. Editing the plan therefore fixes the image without retraining: symbolic verification with resampling gives +5.0 points (p<1e-3), minimal in-place repair +6.0 (p=.02), rewriting only box geometry +10.7 (p<1e-4), and replacing the plan outright +13.3 (p=1e-4). Gains are indifferent to the plan's prose style and to its likelihood under the planner, but not to its geometry. Modular planner-decoder designs are therefore viable, provided the plan is internally consistent: box-text contradictions induce object duplication and identity fusion. We release the plan-fidelity evaluation protocol, all plans, and 12k generated images.

Comments15 pages, 7 figures. Accepted at ECCV 2026 (oral)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑