发表机构
Johns Hopkins University; UC Santa Cruz; Carnegie Mellon University; Rice University(约翰·霍普金斯大学; 加州大学圣克鲁兹分校; 卡内基梅隆大学; 莱斯大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究为视觉模型引入统一表述 RGB In and RGB Out(RINO),将多种视觉信息转为 RGB 图像编辑问题,共享架构参数,基于通用主干无需微调,在多视觉任务上有强大零样本性能,为统一视觉语言系统提供见解。
AI 中文摘要
这项工作为视觉模型引入了一种统一的表述方式,其中除自然图像之外的各种视觉信息形式,如掩码、深度图和其他结构化视觉信号,都被表示为 RGB 图像,而一般视觉任务可转化为常见的 RGB 到 RGB 图像编辑问题。在这种范式下,不同类型的视觉信息在内部与自然图像共享相同的编码、解码架构和参数,使单个模型能够通过统一视觉接口跨任务转移,类似于语言模型处理文本的方式。我们将此表述称为 RGB 输入与 RGB 输出(RINO)。基于通用图像编辑主干且无需特定任务微调,RINO 在密集理解任务(如分割和深度估计,我们将输出统一为 RGB)以及密集条件生成任务(如姿态到图像生成,我们将输入统一为 RGB)上展示了强大且具有竞争力的零样本性能。我们希望这项研究能为通用统一视觉语言系统提供有用见解,即通过共享视觉语言来表达、解释和解决各种视觉任务。代码可从此 https 网址获取。
英文摘要
This work introduces a unified formulation for vision models, where diverse forms of visual information beyond natural images, such as masks, depth maps, and other structured visual signals, are all represented as RGB images, while general visual tasks can be converted into a common RGB-to-RGB image editing problem. In this paradigm, different types of visual information internally share the same encoding and decoding architecture and parameters as natural images, enabling a single model to transfer across tasks through a unified visual interface, in a way analogous to how language models operate over text. We refer to this formulation as RGB In and RGB Out (RINO). Built upon a generic image editing backbone without task-specific fine-tuning, RINO demonstrates robust and competitive zero-shot performance on both dense understanding tasks such as segmentation and depth estimation (where we unify outputs as RGB), and dense-conditioned generation tasks such as pose-to-image generation (where we unify inputs as RGB). We hope this study provides useful insights toward general unified vision-language systems, where diverse visual tasks can be expressed, interpreted, and solved through a shared visual language. Code is available at https://github.com/yangtiming/RINO.