arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

文本到图像模型能否从正确的参考框架中生成图像?

Can Text-to-Image Models Draw from the Right Frame of Reference?

Zheyuan Gu, Ruihang Li, Yong Huang, Yiqian Zhang, XIangzhao Hao, Jiaxin Niu, Jiahao Hu, Zhenyu Zhang

arXiv 2608.03357首次发表:更新:

AI 中文总结

该研究引入FoR-T2I基准,发现现有22个T2I模型在参考框架提示下的布局理解准确率远低于相机视图提示,还提出VLM门控重写方法提升了参考框架下的生成准确率。

AI 中文摘要

空间指令跟随已成为文本到图像(T2I)生成的关键要求。当方向表达在不同参考框架下被解读时,会出现一个常见挑战:例如“在左侧”可能指观察者的图像坐标,也可能指物体的内在朝向,这会导致预期布局不同。现有T2I基准揭示了重要的布局失败,但它们很少区分模型在参考框架与相机视图不同时是否能遵循指定的参考框架。为了缩小这一差距,我们引入FoR-T2I,这是一个用于评估该区别的基准,包含1200个由受控空间布局构建的提示对。在每一对中,相机视图(Cam)提示以相机视图陈述目标关系,而参考框架(FoR)提示则通过定向锚定物体描述相同的目标位置。在22个闭源和开源T2I模型中,FoR提示的平均最终准确率比匹配的Cam提示低41.8%;即使是表现最好的模型,FoR准确率也仅达到44.3%。这表明,当相同布局通过物体朝向而非直接以图像坐标描述时,当前模型的表现更差。我们进一步按关系类型和相机视图分析了这一差距,比较了几种无训练的提示和基于反馈的缓解策略,并提出了一种VLM门控重写方法,该方法利用视觉反馈选择重写后的提示,在相同生成预算下将平均FoR准确率从25.0%提升至29.2%。

英文摘要

Spatial instruction following has become a crucial requirement for text-to-image (T2I) generation. A common challenge arises when directional expressions are interpreted under different frames of reference. For example, ``the left of'' may refer to the viewer's image coordinates or to the intrinsic orientation of an object, leading to different expected layouts. Existing T2I benchmarks reveal important layout failures, yet they rarely isolate whether models can follow a specified frame of reference when it differs from camera view. To mitigate this gap, we introduce FoR-T2I, a benchmark for evaluating this distinction with 1,200 prompt pairs built from controlled spatial layouts. In each pair, the camera-view (Cam) prompt states the target relation in camera view, while the frame-of-reference (FoR) prompt describes the same target placement through an oriented anchor object. Across 22 closed-source and open-source T2I models, mean final accuracy is 41.8\% lower on FoR prompts than on matched Cam prompts; even the best-performing model achieves only 44.3\% FoR accuracy. This suggests that current models struggle more when the same layout is described through an object's orientation rather than directly in image coordinates. We further analyze this gap by relation type and camera view, compare several training-free prompting and feedback-based mitigation strategies, and propose a VLM-gated rewriting approach that selects rewritten prompts using visual feedback, improving average FoR accuracy from 25.0\% to 29.2\% under the same generation budget.

Comments9 pages, 3 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑