AI 中文总结
本研究通过伪造任务实测ChatGPT Images 2.5,发现其Flare和Sunburst模型在减少收据文本改动上优于GPT-Image-2,但目标字段正确性无提升,且保真度改进有限,强调需任务特定评估宣称能力。
AI 中文摘要
我们评估了ChatGPT Images 2.5宣称的改进是否能在具有预定答案的伪造任务中转化为更好的性能。我们将它的Flare和Sunburst API模型与同一周内重新运行的GPT-Image-2进行比较,使用收据字段编辑、重复编辑、产品植入和细打印渲染等任务。在图像配准后,Flare和Sunburst对周围收据文本的OCR检测变化较少(分别为31.7%和31.2%,而两个GPT-Image-2基线均为44.2%),主要出现在CORD收据上,但在目标字段正确性上没有检测到明显改进。Flare在CORD收据上保留了更少的早期编辑,而照片编辑序列在模型之间几乎没有区分度。产品代码在Images 2.5下更常清晰可读,同时产品植入更大;分析并未确立独立于尺寸的保真度提升。细打印的改进在OCR可靠性限制以下仍未解决。两代模型的拒绝(弃权(不执行))情况罕见,且定位能力较弱。在固定检测阈值下,Community Forensics在单元格平均上标记了68.6%的受控Images 2.5图像,而在线发布的自我报告图像为35.9%。这些结果促使对宣称的能力和防御措施进行任务特定评估,并明确自动检查能确立的局限性。
英文摘要
OpenAI released ChatGPT Images 2.5 on 8 September 2026, advertising more precise local edits, better consistency across edits, more faithful reference products and sharper detail. We evaluate these claims on four forgery tasks with answers fixed in advance: receipt-field alteration, repeated editing, product placement and small-print rendering. GPT-Image-2 provides same-week baselines at a cheaper and a more expensive tier. A limited improvement appears in receipt editing. After alignment, OCR detects changes to surrounding text in 31.7% of Flare outputs, against 44.2% for the cheaper baseline. This gain is concentrated on CORD receipts and sensitive to shifts of a pixel or less; the forged value itself is no more often correct. Repeated editing and fine print show no measurable gain. Product codes become more legible mainly because Images 2.5 draws the product larger. Defence outcomes change little: localisation remains weak for both generations. A detector that flags 68.6% of controlled benchmark images flags only 35.9% of images posted online. Advertised improvements therefore transfer unevenly to the tested forgery capabilities, while substantial detection limitations remain.
Comments27 pages, 6 figures, 16 tables