图像生成器是零样本感知器吗?一项严格的评估
Are Image Generators Zero-Shot Perceivers? A Rigorous Evaluation
浏览论文内容
中文总结 AI 辅助
本研究提出ProbeGen基准,系统评估图像生成器在零样本视觉感知任务中的能力,发现其具备一定感知能力但专业模型在分布内更优,生成模型在分布偏移和语义推理上更鲁棒。
中文摘要 AI 辅助
近期的工作,如Vision Banana,表明轻量级指令微调可以使图像生成器在多个视觉感知任务上达到最先进的性能。受此观点启发,我们探究在零样本设置下,图像生成器在公开视觉感知基准上能走多远。我们引入了ProbeGen,一个用于零样本生成感知的基准,它将单目深度估计、指代/推理分割和物体计数转化为通过文本提示指定的条件生成任务,并总共比较了20个模型——包括专有和开源权重的图像生成器、专业感知模型以及多模态大语言模型(MLLMs)——跨越11个已发表的基准。我们观察到,预训练的图像生成器表现出可测量的零样本感知能力,但存在明显的权衡:专业模型在分布内准确性和效率方面仍然更强,而生成模型在分布偏移下通常更鲁棒,并且在组合语义推理方面表现更好。我们希望这项研究有助于将零样本生成感知确立为一个有意义的研究方向,并为未来视觉生成与理解交叉领域的工作提供有用的基础。
英文摘要
Recent work, such as Vision Banana, shows that lightweight instruction tuning can enable an image generator to achieve state-of-the-art performance across multiple visual perception tasks. Motivated by this perspective, we ask how far image generators can go on public visual perception benchmarks in a zero-shot setting. We introduce ProbeGen, a benchmark for zero-shot generative perception that casts monocular depth estimation, referring/reasoning segmentation, and object counting as conditional generation tasks specified through text prompts, and compares 20 models in total---including proprietary and open-weight image generators, specialist perception models, and MLLMs---across 11 published benchmarks. We observe that pretrained image generators show measurable zero-shot perceptual competence, but with a clear trade-off: specialist models remain stronger for in-distribution accuracy and efficiency, while generative models are often more robust under distribution shift and better at compositional semantic reasoning. We hope this study helps establish zero-shot generative perception as a meaningful research direction and provides a useful foundation for future work at the intersection of visual generation and understanding.
发表机构
- Shanghai Jiao Tong University(上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。