arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

总是10:10:参考图像打破仅提示时钟的偏差

It's Always 10:10: Reference Images Break a Bias That Prompts Only Dent

Luca Cazzaniga

arXiv 2610.11320首次发表:更新:

AI 中文总结

该研究发现文本到图像模型存在时钟显示10:10的偏差,在52种模型上测试三种方法,添加绘制参考表盘可将时钟完全准确率提升至75%,几乎消除10:10偏差。

AI 中文摘要

文本到图像模型似乎会复制其训练所用照片的习惯,模拟时钟是一个极端案例:在广告中,手表几乎总是显示10:10,即便要求显示其他时间,生成的时钟也会回到10:10。我们在Magnific平台上的52种模型上测量了这种偏差,并测试了三种克服方法,还在Higgsfield平台上进行了重复实验。每张图像显示三个相同的时钟,需分别显示2:35、6:50和11:20。物体描述固定,仅时间要求变化:无时间要求(A)、数字形式的时间(B)、指针相对于表盘数字的位置描述(C),或相同描述加绘制的参考表盘(D)。两名AI读取器从编码副本中盲读全部1799张图像,第三名读取器及作者解决分歧(表盘级别一致性为96.0%和97.3%)。无时间要求时,67%的图像中三个时钟均为10:10。在20种当前模型中,数字形式时间要求下34%的图像三个时钟均正确,指针文字描述下为30%,参考表盘(D)下为75%(D-B差值为+37个百分点,95%置信区间+28至+45);我们未发现指针文字描述优于数字形式的证据(C-B差值为-4个百分点,置信区间-10至+1)。两个平台共有的12种模型的重复实验结果一致(B为54%,C为50%,D为81%)。写入时间可降低偏差,但20种当前模型中仍有三分之二的图像至少有一个时钟错误;添加绘制参考表盘可将完全准确率提升至四分之三,几乎消除了全为10:10的图像。我们发布了所有图像、提示、原始读数及可重新计算所有结果的脚本。

英文摘要

Text-to-image models appear to reproduce the habits of the photographs they learned from. Analog clocks are an extreme case: in advertising, watches almost always show 10:10, and generated clocks return to 10:10 even when another time is requested. We measure this bias and test three ways of overcoming it on 52 models available on the Magnific platform, with a replication on Higgsfield. Every image shows three identical clocks that must show 2:35, 6:50 and 11:20. The description of the object is fixed and only the request about the time changes: no time (A), the time in digits (B), the hand positions described by construction relative to the dial numerals (C), or the same description plus a drawn reference dial (D). Two AI readers read all 1,799 images blind from coded copies, with a third reader and the author settling disagreements (dial-level agreement 96.0% and 97.3%). With no time requested, 67% of the images have all three clocks at 10:10. On the 20 current models, all three clocks are correct in 34% of the images with digits, 30% with the hands described in words and 75% with the reference dial (D-B: +37 points, 95% CI +28 to +45); we found no evidence that describing the hands in words beats the digits (C-B: -4 points, CI -10 to +1). The replication on the 12 models shared by both platforms gives the same picture (B 54%, C 50%, D 81%). Writing the time reduces the bias but leaves two thirds of the images of the 20 current models with at least one wrong clock; adding a drawn reference raises full accuracy to three quarters and almost eliminates images entirely at 10:10. We release all images, prompts, raw readings and a script that recomputes every result.

Comments16 pages, 5 figures, 6 tables. Data and code: doi:10.5281/zenodo.23224681

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑