发表机构
Data Science & Artificial Intelligence Research Institute, China Unicom; Unicom Data Intelligence, China Unicom; China Unicom Group Co.,Ltd(中国联通数据科学与人工智能研究院; 中国联通数据智能; 中国联通集团有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对T2I模型评估的问题,提出QC-T2I-Bench框架,结合DSG实现细粒度诊断,评估显示多能力组件联合完成率随能力数增加下降,成本感知路由器可在减少算力的同时达到ERNIE的估计精度。
AI 中文摘要
现代文本到图像(T2I)模型常具有相近的总得分但优势不同,导致实际选择困难。细粒度基准将提示分解为问题,但常返回提示级得分与固定类别,削弱归因性且忽略复杂性;相关需求也被单独评分或合并为总得分,模糊了基础失败与组合失败的区别。我们提出QC-T2I-Bench,这是一个以问题为中心的框架,将开放式提示转换为可归因的原子问题,并使用戴维森场景图(DSGs)组织它们的依赖关系。我们采用层次约束的问题聚合,在前提失败后排除下游问题,防止简单和复杂提示获得相同的总权重。随后,我们利用DSG结构测量提示内的联合成功,比较不同提示间的重复实体,将基础实现失败与附加要求下的失败区分开。我们在英文和中文提示上评估多个开源T2I模型,所得问题级证据支持可靠排名与细粒度诊断:具备两种能力的组件联合完成率为80.7%,具备七种或更多能力的组件联合完成率降至37.2%。最后,我们复用相同记录进行无训练路由,我们的成本感知路由器以少21.3%的GPU秒/MP,达到与ERNIE相同的89.51点估计值。
英文摘要
Modern text-to-image (T2I) models often have similar total scores but different strengths, making practical selection difficult. Fine-grained benchmarks decompose prompts into questions, yet often return them to prompt scores and fixed categories, weakening attribution and ignoring complexity. Related requirements are also scored separately or as one total, obscuring basic versus compositional failure. We present QC-T2I-Bench, a question-centric framework that converts open prompts into attributed atomic questions and organizes their dependencies with Davidsonian Scene Graphs (DSGs). We use hierarchy-constrained question aggregation to exclude downstream questions after a prerequisite fails and to prevent simple and complex prompts from receiving the same total weight. We then use the DSG structure to measure joint success within prompts and compare repeated entities across prompts, separating basic realization failures from failures under additional requirements. We evaluate multiple open-source T2I models on English and Chinese prompts. The resulting question-level evidence supports reliable ranking and fine-grained diagnosis: joint completion falls from 80.7\% for components with two capabilities to 37.2\% for those with seven or more. Finally, we reuse the same records for training-free routing; our cost-aware router matches ERNIE's 89.51-point estimate with 21.3\% less GPU-s/MP.