arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PixelUp:用于细粒度视觉任务的零样本语义特征上采样

PixelUp: Zero-Shot Semantic Feature Upsampling for Fine-Grained Vision Tasks

Deepank Singh, Anurag Nihal, Vedhus Hoskere

arXiv 2608.02792首次发表:更新:

AI 中文总结

PixelUp是一种零样本、与视觉基础模型无关的语义特征上采样器,通过多尺度语义引导的窗口交叉注意力架构实现细粒度视觉任务的最优性能,可提升语义分割、深度估计等任务的精度。

AI 中文摘要

自监督视觉基础模型(Vision Foundation Models, VFMs)凭借其强大且可迁移的视觉表征,已成为下游任务不可或缺的骨干模型。然而,当需要准确的细粒度预测时,它们的补丁标记级(patch-token-level)特征对于语义分割、深度估计等密集预测任务而言往往过于粗糙。现有特征上采样方法虽已被开发用于恢复像素级细节,但仍存在局限:可学习上采样器通常针对特定编码器设计,若换用不同编码器则必须重新训练;基于浅层像素编码器的图像引导方法常引入纹理伪影,且缺乏准确下游预测所需的语义引导。本文提出PixelUp,这是一种与VFMs无关的零样本上采样器,其通过由多尺度语义特征引导的、从粗到细的窗口交叉注意力架构链实现语义感知。我们的实验表明,PixelUp的性能优于特定于VFMs和与VFMs无关的上采样器,在密集预测任务上达到了当前最佳性能:在跨多个VFMs的NYUv2深度估计任务中,语义分割的平均mIoU提升了1.2,δ₁提升了0.25;PixelUp还进一步提升了无训练的开放词汇语义分割和无监督语义分割,平均mIoU分别提升了1.3和0.5。代码可在该https URL获取。

英文摘要

Self-supervised Vision Foundation Models (VFMs) have become essential backbones for downstream tasks due to their strong and transferable visual representations. However, their patch-token-level features are often too coarse for dense prediction tasks such as semantic segmentation and depth estimation when accurate fine-grained predictions are required. Feature upsampling methods have been developed to recover pixel-level detail but still face limitations. Learnable upsamplers are often designed for a specific encoders and must be retrained for different encoders. Image-guided methods that use shallow pixel encoders often introduce textural artifacts and lack the semantic guidance needed for accurate downstream predictions. We introduce PixelUp, a zero-shot VFM-agnostic upsampler achieving semantic awareness through a coarse-to-fine chain of windowed cross-attention architecture guided by multi-scale semantic features. We demonstrate that PixelUp outperforms both VFM-specific and VFM-agnostic upsamplers, achieving state-of-the-art performance on dense prediction tasks with an average improvement of +1.2 mIoU on semantic segmentation and +0.25 $δ_1$, on NYUv2 depth estimation across VFMs. PixelUp further improves training-free open-vocabulary and unsupervised semantic segmentation by an average of +1.3 mIoU and +0.5 mIoU, respectively. Code available at https://pixelup-project.vercel.app/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑