OmniTaskonomy:视觉生成何时提升视觉理解?
OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?
- University of California, Berkeley(加州大学伯克利分校)
- Duke University(杜克大学)
- Carnegie Mellon University(卡内基梅隆大学)
- University of Washington(华盛顿大学)
- Elorian
- Impossible Research
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究通过OmniTaskonomy分类体系系统探究视觉生成监督对视觉理解的影响,发现特定生成任务能选择性提升对应理解能力,且梯度对齐程度与迁移增益正相关,为视觉理解训练提供指导。
AI中文摘要:
训练一个模型来生成视觉内容可以鼓励其学习与几何、空间关系和物体性相关的丰富感知能力;然而,这对视觉理解的好处仍不清楚。我们提出疑问:视觉生成监督在何时以及如何提升视觉理解?我们研究了成对的图像到图像(I2I)生成任务和图像到文本(I2T)理解任务,这些任务以不同的输出模态表达相同的底层问题。我们发现,在正确的配方下,I2I训练能提升下游I2T性能,且随着I2I训练数据量的增加,提升幅度更大。接下来,我们探讨哪些生成任务有益于哪些理解能力。为了研究配对任务之外的迁移,我们引入了OmniTaskonomy,一个统一的分类体系,涵盖19个I2I生成任务和25个I2T理解能力。由此产生的迁移图揭示了选择性的、任务依赖的益处。一些迁移遵循直观的对应关系,例如,深度预测提升度量3D推理,物体指向提升计数,拼图重建提升2D排序。有趣的是,我们还发现了令人惊讶的联系:2.5D分割提升类别识别,Z深度预测提升定位。为了探究这些模式,我们分析了生成任务与理解任务之间的梯度对齐,发现更强的对齐与更大的下游迁移增益相关。总之,我们的结果凸显了视觉生成作为视觉理解丰富监督来源的潜力,并为通过正确的训练课程和任务选择来释放其益处提供了路线图。项目页面:此https URL。
英文摘要:
Training a model to generate visual content can encourage it to learn rich perceptual capabilities related to geometry, spatial relationships, and objectness; yet, its benefits for visual understanding remain unclear. We ask: when and how does visual generation supervision improve visual understanding? We study controlled pairs of image-to-image (I2I) generation and image-to-text (I2T) understanding tasks that express the same underlying problem in different output modalities. We find that under the correct recipe, I2I training improves downstream I2T performance, with larger gains as the amount of I2I training data increases. We next ask which generation tasks benefit which understanding capabilities. To study transfer beyond paired tasks, we introduce OmniTaskonomy, a unified taxonomy spanning 19 I2I generation tasks and 25 I2T understanding capabilities. The resulting transfer map reveals selective, task-dependent benefits. Some follow intuitive correspondences, e.g., depth prediction improving metric 3D reasoning, object pointing improving counting, and jigsaw reconstruction improving 2D ordering. Interestingly, we also uncover surprising connections: 2.5D segmentation improving category recognition and Z-depth prediction improving localization. To probe these patterns, we analyze gradient alignment between generation and understanding tasks and find that stronger alignment is associated with larger downstream transfer gains. Together, our results highlight visual generation as a rich source of supervision for visual understanding and provide a roadmap for unlocking its benefits through the right training curriculum and task selection. Project page: https://omni-taskonomy.github.io/.