ChatGPT Images 2.5 实战观察:发布期数据集与检测器评估
ChatGPT Images 2.5 in the Wild: A Launch-Period Dataset and Detector Evaluation
浏览论文内容
中文总结 AI 辅助
针对ChatGPT Images 2.5发布期,构建含3,478张图像的冻结数据集,评估六个检测器性能,发现标记率差异大且与召回率无直接关联,支持版本归属分析。
中文摘要 AI 辅助
图像工具可以在保留其公开名称的同时更换底层生成器,这使得从在线帖子进行版本归属变得模糊不清。我们在 ChatGPT Images 2.5 发布后研究了这一问题。我们的冻结数据集包含来自 8 个来源、2,440 个帖子的 3,478 张图像。记录的发布时间落在公告发布后的前 51.1 小时内。数据集记录了三个归属层级,并在图像形式过滤和定向审查后保留了独立图像。标题声明和宿主记录提供了准入证据,但并非独立验证的生成器身份。观察到的内容分布取决于来源混合:NightCafe 提供了 39.0% 的图像,但贡献了 77.0% 的 CLIP 分配的幻想场景。随后,我们评估了六个冻结检测器,其阈值校准为在参考照片上 5% 的标记率。数据集的标记率范围为 3.7% 至 56.4%,比 GenImage 召回率低 42 至 81 个百分点。保留的 artworks 假阳性率范围为 1.5% 至 96.5%,因此较高的数据集标记率本身并不能证明更好的检测性能。一项探索性的仅限 X 平台的比较(与我们的四月数据集相比)发现,在固定阈值、聚类后自举区间下,Effort 的九月标记率更高,而 DoU 存在提示性差异。归属、内容和处理差异阻止了对这些对比的因果解释。该数据集支持在产品过渡期间对报告模型使用情况的分析,并保留了来源和归属证据以供解释。
英文摘要
An image tool can change its underlying generator while retaining its public name, making version attribution from online posts ambiguous. We study this problem after the ChatGPT Images 2.5 launch. Our frozen collection contains 3,478 images from 2,440 posts across 8 sources. Recorded posting times fall within the first 51.1 hours after the announcement. It records three attribution tiers and retains standalone images after image-form filtering and targeted review. Caption claims and host records provide admission evidence, not independently verified generator identity. The observed content profile depends on the source mixture: NightCafe supplies 39.0% of images but 77.0% of CLIP-assigned fantasy scenes. We then evaluate six frozen detectors at thresholds calibrated to a 5% flag rate on reference photographs. Collection flag rates range from 3.7 to 56.4%, falling 42-81 percentage points below GenImage recall. Held-out artwork false-positive rates range from 1.5 to 96.5%, so a higher collection flag rate does not by itself establish better detection. An exploratory X-only comparison with our April collection finds a higher September flag rate for Effort, and a suggestive difference for DoU, under fixed-threshold post-clustered bootstrap intervals. Attribution, content and processing differences prevent a causal interpretation of these contrasts. The collection supports analysis of reported model use during a product transition, with source and attribution evidence retained for interpretation. The collection is released at https://scam.ai/research.