arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.28206cs.CVcs.DB

NumBench:诊断文本到图像模型中的计数失败问题

NumBench: Diagnosing Counting Failures in Text-to-Image Models

  • Indian Institute of Technology Jodhpur(焦特布尔印度理工学院)
  • Shiv Nadar University(希夫·纳达尔大学)

机构由 AI 辅助整理,请以论文原文为准。

Sandeep Wadhwa, Mayank Vatsa, Richa Singh, Parrva Chirag Shah, Prakhar Galriya

AI总结:

该研究推出NumBench基准数据集,构建过程模型与置信度加权数值精确率指标,评估发现文本到图像模型计数性能随请求数增加而下降,50个以上时表现薄弱,且方法可跨模板迁移。

AI中文摘要:

文本到图像(T2I)模型经常生成数量错误的对象,然而现有的基准数据集规模过小或控制力度不足,无法解释出现该问题的原因。我们推出NumBench,这是一个包含640000个提示的基准数据集,涵盖1600个类别以及1到100的计数范围。其因子设计会调整对象组成、空间引导和外观条件,同时平衡计数与类别曝光度。我们还构建了一个过程模型,其中请求的实例会竞争有限数量的可解析图像区域。该模型预测,在低占用率下会出现接近二次方的碰撞缺失,并且展示了如何通过协调放置来减少这种缺失。为实现可扩展评估,我们提出了置信度加权数值精确率(Confidence-Weighted Numeric Precision Score,简称cwnps),该指标整合了三个经过校准的检测器,并对不确定的候选结果进行折扣处理。在对五个商业系统、两个开放模型以及两个专用计数方法的测试中,性能随请求计数的增加而急剧下降;所有被评估的方法在超过50个对象时表现都很薄弱。计数范围是已测量到的最大影响因素,其次是布局和组成。在引导布局中,网格引导的效果最强,这与协调放置的预测一致,不过该分析并未证实碰撞是唯一的原因。一项包含14400张图像的人类研究支持将自动评估应用到计数50的场景,而在243个自然语言提示上的结果表明,该基准的评估方法可以在NumBench模板之外进行迁移。

英文摘要:

Text-to-image (T2I) models often generate the wrong number of objects, yet existing benchmarks are too small or weakly controlled to explain why. We introduce \textbf{NumBench}, a benchmark of 640{,}000 prompts spanning 1{,}600 categories and counts from 1 to 100. Its factorial design varies object composition, spatial guidance, and appearance conditions while balancing counts and category exposure. We also develop a process model in which requested instances compete for a finite set of resolvable image regions. The model predicts a near-quadratic collision deficit at low occupancy and shows how coordinated placement reduces it. For scalable evaluation, we propose the Confidence-Weighted Numeric Precision Score (\cwnps), which aggregates three calibrated detectors and discounts uncertain proposals. Across five commercial systems, two open models, and two specialized counting methods, performance declines sharply with requested count; all evaluated methods are weak above 50 objects. Count range has the largest measured effect, followed by layout and composition. Grid guidance is strongest among guided layouts, consistent with the coordination prediction, although the analysis does not establish collision as the sole cause. A 14{,}400-image human study supports automated evaluation through count 50, while results on 243 natural-language prompts show transfer beyond NumBench templates.

补充信息

↑