发表机构
Nokia; University of Georgia; The Hong Kong Polytechnic University(诺基亚; 佐治亚大学; 香港理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出代码-图像推理范式,通过无训练反思循环让模型自我演进技能,在CwI-Bench上显著提升多模态模型视觉任务准确率,技能可跨规模与任务家族迁移。
AI 中文摘要
多模态模型在解决视觉任务时越来越多地借助工具(裁剪、缩放、旋转、调亮),这一范式被称为“图像思考”。核心挑战在于感知:工具主要用于呈现视觉证据,对证据的推理仍停留在语言层面,且多数目标是人类原则上可通过检查确定的。然而,一些视觉问题的瓶颈并非感知:要得出答案,需对像素执行多步视觉算法。在这类问题上,模型常能立刻说出正确算法却仍答错,因为语言可描述算法却无法运行算法。代码-图像突破了这一局限:仅提供Python解释器,模型必须用代码实现真正的视觉算法来解决任务,程序本身即成为推理。此时瓶颈从执行代码转移到决定实现何种算法。因此,我们让模型自我学习:一个无训练的反思循环研究自身失败的程序,对照建设性真值测试修复方案,保留存活的内容作为可迁移技能。在我们的代码-图像基准(CwI-Bench)上,由隐藏视觉计算构成的30个任务家族具有不相交的训练和评估拆分,即使是GPT-5.6-luna,无工具的思维链准确率仍低于30%;仅提供裸解释器时,准确率达43%;经自身可执行反思演进的技能加持后,准确率达67%。开源27B模型也遵循相同提升路径(9%→33%→56%),且这些技能为纯文本,可跨规模和任务家族迁移。当代码承载推理时,调试代码即调试推理。
英文摘要
Multimodal models increasingly reach for tools when solving visual tasks (crop, zoom, rotate, brighten), a paradigm known as thinking-with-images. The central challenge is one of perception: tools mostly serve to expose visual evidence, reasoning over that evidence stays in language, and most targets are ones a human could in principle determine by inspection. Some visual questions, however, are not bottlenecked by perception: recovering their answers requires executing a multi-step visual algorithm over the pixels. On such questions a model often names the correct algorithm at once yet still answers wrong, because language can describe an algorithm without being able to run one. Code-with-Image crosses that line: given nothing but a Python interpreter, the model must implement a genuine visual algorithm in code to solve the task; the program itself becomes the reasoning. The bottleneck then shifts from executing code to deciding which algorithm to implement. So we let the model teach itself: a training-free reflection loop studies its own failed programs, tests repairs against constructive ground truth, and keeps what survives as portable skills. On our Code-with-Image Bench (CwI-Bench), thirty task families induced by hidden visual computations with disjoint learning and evaluation splits, even GPT-5.6-luna stays below 30% with tool-free chain of thought; given a bare interpreter it reaches 43%, and with skills evolved through its own executable reflection, 67%. The open 27B model climbs the same ladder (9% $\rightarrow$ 33% $\rightarrow$ 56%), and the skills are plain text, transferable across scales and families. When code carries the reasoning, debugging code becomes debugging reasoning.
Comments37 pages