以对象为中心的LeJEPA
Mask-supervised Object-centric Representation Learning with LeJEPA
浏览论文内容
中文总结 AI 辅助
提出以对象为中心的LeJEPA方法,通过使用SAM提供的对象掩码,在对象级别对齐表示,避免场景级自监督方法的数据低效和不稳定性,在跟踪、分类、分割和重识别任务上优于图像级LeJEPA。
中文摘要 AI 辅助
使用LeJEPA训练的图像编码器可以为下游任务提供强大的特征,但与其他图像级自监督方法一样,通常需要大型训练数据集。在对象级别而非整个场景对齐表示有望提高数据效率,但以完全自监督的方式实现这一点,即有效联合分割场景并表示其对象,是不稳定的:两者陷入循环依赖,分割需要有意义的表示,而有意义的表示需要一致的分割。我们通过在训练期间将对象掩码作为给定条件来规避这种不稳定性,使用廉价的现成SAM提议。我们将LeJEPA——其分布性抗塌缩目标自然地适用于从整个图像到可变大小的对象集——扩展到对齐对象中心表示而非整个图像。一个额外的实例分离损失,将同一场景中的其他对象视为负样本,进一步提升了下游性能。在两个模型规模和COCO的10-100%数据上,对象级LeJEPA在跟踪(DAVIS)、分类(ImageNet-1k)、分割(ADE20k)和重识别(NAVI)任务上优于图像级LeJEPA。
英文摘要
Self-supervised image encoders deliver strong features for downstream tasks but need many images for training. A natural remedy to counter this is to make each image count for more. A scene contains many objects, and given masks from human annotators or an off-the-shelf segmentation model, pre-training can focus on aligning per-object rather than image-wide representations, extracting more signal from every image. Existing mask-supervised methods do this through reconstruction or contrastive losses that leverage negative objects. We instead use two separate projection spaces for the alignment. In a \emph{semantic space}, per-object representations from different views are aligned. To avoid collapse, instead of using negative objects, which requires category definitions, we extend the negative-free LeJEPA objective and show that its distributional anti-collapse regularizer ports naturally from whole images to the variable-sized set of objects in a scene. In an \emph{instance space}, a contrastive loss separates per-object representations from their context and co-occurring instances, including those of the same category. To separate object representations from their context, we copy objects and paste them into other contexts, where each pasted copy serves as an additional view of the original object. Trained on COCO with ground-truth masks, our method outperforms image-level and mask-guided baselines on tracking (DAVIS), classification (ImageNet-1k) and re-identification (NAVI), matches the best of them on semantic segmentation (ADE20k) and keeps its lead over image-level LeJEPA and a supervision-matched alternative on COCO fractions down to 256 images.
发表机构
- Biomedical Image Computing Group, ETH Zurich(苏黎世联邦理工学院生物医学图像计算组)
机构由 AI 辅助整理,请以论文原文为准。