发表机构
The University of Tokyo; National Institute of Advanced Industrial Science and Technology (AIST); University of Oxford(东京大学; 独立行政法人产业技术综合研究所; 牛津大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出WISRD基准测试图像编辑模型的图像空间规则发现能力,发现Nano Banana Pro表现最优,当前模型可部分依赖图像内指令,且该模型在4×4数独等任务上有一定性能。
AI 中文摘要
图像编辑模型能否在图像空间中发现视觉规则并端到端完成问题解决?我们以人类工作表测试(如智商测试)的思路来解决这一问题,所用问题要求模型读取基于图像的指令、识别问题、推断答案、将答案绑定到正确位置、控制输出数量、抑制不必要的编辑、并保留输入内容和格式。我们推出WISRD(Worksheet Image-Space Rule Discovery,工作表图像空间规则发现)基准,包含8种信息条件下的11项核心任务,涵盖局部标记、填充、复制、计数和无编辑抑制,另有4项补充推理压力探针,用于多步空间操作、抽象模式推理、逻辑推理及基于约束的问题解决。我们确定三项关键发现:(i)在评估的前沿图像编辑模型中,Nano Banana Pro得分最高;在共享的V0--V3无参考子集上,Auto-Strict代理通过率为:Nano Banana Pro 48.7%、Qwen-Image-Edit 13.4%、FLUX.2 Klein 4B API 11.5%、FLUX.2 Klein 4B open-weight 11.3%、InstructPix2Pix 0.0%。(ii)分析显示,即使外部提示缺失或仅为通用内容,当前图像编辑模型仍可部分依赖渲染的图像内指令。(iii)在小型补充诊断任务中,Nano Banana Pro在4×4数独任务上准确率达70.0%,在图像空间的公开RAVEN模式发现项上准确率为22.9%。
英文摘要
Can image-editing models discover visual rules in image space and complete problem-solving end-to-end? We tackle this question in the spirit of a human worksheet test (e.g., an IQ test), using problems that require models to read image-based instructions, recognize the problem, infer the answer, bind it to the correct destination, control output count, suppress unnecessary edits, and preserve the input and format. We introduce WISRD, a Worksheet Image-Space Rule Discovery benchmark with 11 core tasks under eight information conditions, spanning localized marking, filling, copying, counting, and no-edit suppression, together with four supplementary reasoning-stress probes for multi-step spatial manipulation, abstract pattern reasoning, logical inference, and constraint-based problem solving. We identify three key findings as follows. (i) Among the frontier image-editing models evaluated, Nano Banana Pro achieves the highest score. On the shared V0--V3 no-reference subset, the Auto-Strict proxy pass rates are 48.7% for Nano Banana Pro, 13.4\% for Qwen-Image-Edit, 11.5% for FLUX.2 Klein 4B API, 11.3% for FLUX.2 Klein 4B open-weight, and 0.0% for InstructPix2Pix. (ii) Analysis reveals that current image-editing models can partially rely on rendered in-image instructions even when the external prompt is absent or merely generic. (iii) In small supplementary diagnostics, Nano Banana Pro achieves 70.0% on 4-by-4 Sudoku and 22.9% on public RAVEN pattern-discovery items in image space.
Comments20 pages, 5 figures