TangPoetryBench:面向诗歌到图像生成的多维度基准与基于评分规则的评估器
TangPoetryBench: A Multi-Dimensional Benchmark and Rubric-Conditioned Evaluator for Poetry-to-Image Generation
中文总结 AI 辅助
该研究推出TangPoetryBench多维度基准与PAE评估器,解决T2I模型生成诗歌插图的评估难题,PAE性能接近Claude且可泛化,相关资源已公开。
中文摘要 AI 辅助
文本到图像(T2I)模型越来越多地被要求呈现文学和文化内容,但我们无法衡量图像对诗歌含义的呈现效果。该任务是多方面的:一幅好的插图必须视觉效果良好,忠实于诗歌的意象和场景,符合文化和风格要求,无多余文字,且契合诗歌的情感;其最核心的要求,即意象尤其是隐含的情感,从未在文字中明确表述。现有指标(CLIPScore、BLIPScore、VQAScore)奖励字面的文本-图像对应关系,因此无法判断一幅插图是否成功,更无法说明原因,甚至无法区分最佳模型与最差模型。我们推出TangPoetryBench,这是一个包含1280张图像的多维度基准(320首中国唐代古典诗歌 × 4个最先进的T2I模型),涵盖十个维度的质量控制人工标注。通过分析该数据,我们揭示了当前T2I模型的共性优势与模型特定的优缺点,包括其唤起诗歌隐含情感的能力。我们进一步推出PoemAutoEvaluator(PAE),这是一个开源的、基于评分规则的评估器,其性能可与强大的专有评判模型Claude相媲美,能泛化到未见过的生成器和第二种诗歌传统(宋词),并使基准能够在无需新人工标注的情况下扩展到新图像。我们发布该基准、标注和评估器。
英文摘要
Text-to-image (T2I) models are increasingly asked to illustrate literary and cultural content, yet we cannot measure how well an image renders the meaning of a poem. The task is many-sided: a good illustration must be visually sound, faithful to the poem's imagery and scene, culturally and stylistically apt, free of spurious text, and true to its emotion, and its deepest requirements, imagery and especially implicit emotion, are never stated in the words. Existing metrics (CLIPScore, BLIPScore, VQAScore) reward literal text-image correspondence and so cannot tell whether an illustration succeeds, let alone why, or even separate the best model from the worst. We introduce TangPoetryBench, a multi-dimensional benchmark of 1,280 images (320 classical Chinese Tang poems x 4 state-of-the-art T2I models) with quality-controlled human annotations across ten dimensions. Analyzing this data, we reveal the shared and model-specific strengths and weaknesses of current T2I models, including their ability to evoke a poem's implicit emotion. We further introduce PoemAutoEvaluator (PAE), an open, rubric-conditioned evaluator that reaches parity with a strong proprietary judge (Claude), generalizes to an unseen generator and a second poetic tradition (Song Ci), and lets the benchmark scale to new images without fresh human annotation. We release the benchmark, annotations, and evaluator.