发表机构
School of Electronic Information and Electrical Engineering, Shanghai Jiao Tong University; College of Computing and Data Science, Nanyang Technological University; Pengcheng Laboratory; Ant Group(上海交通大学电子信息与电气工程学院; 南洋理工大学计算与数据科学学院; 鹏城实验室; 蚂蚁集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对ViT特征编码忽略二维补丁网格结构的问题,提出双路径视觉令牌编解码器VTC,在DINOv2、SAM3的分类等任务上,码率降低15.7-37.4倍且性能达未压缩的90%,表现优于基线。
AI 中文摘要
大型视觉基础模型的分布式部署通常会划分ViT主干网络,并在计算节点间交换中间令牌特征,因此在带宽与计算资源约束下,高效的特征压缩至关重要。现有ViT特征编解码器通常将异构的全局令牌与补丁令牌展平为L×C的伪图像,导致熵模型主要捕捉序列轴依赖,却忽略了其原生的二维补丁网格结构。本文中,我们发现ViT补丁令牌在原始网格上保留着强烈的局部空间相关性。为利用这一结构先验,我们提出了视觉令牌编解码器(Visual Token Codec, VTC),这是一种双路径学习型编解码器,将全局令牌与补丁令牌分离到专用编码路径:全局令牌采用轻量级因子化先验进行压缩,补丁令牌则采用空间-通道上下文熵模型在补丁令牌网格上编码。为支持中间层压缩与实际码率自适应,VTC进一步在后续ViT块后引入特征匹配监督,并在单个编解码器内集成可变速率模块。在DINOv2与SAM3上的实验表明,VTC在分类、分割、检测任务上均显著优于代表性ViT特征编码基线;在达到未压缩特征90%性能时,VTC在这些任务上将码率降低15.7倍至37.4倍。我们还针对面向传输与存储的实际部署场景,提供了中间层的码率-效用分析。
英文摘要
Distributed deployment of large vision foundation models often partitions a ViT backbone and exchanges intermediate token features between computing nodes, making efficient feature compression critical under bandwidth and computation constraints. Existing ViT feature codecs typically flatten heterogeneous global and patch tokens into an L x C pseudo image, causing entropy models to mainly capture sequence-axis dependencies while overlooking the native two-dimensional patch-grid structure. In this paper, we show that ViT patch tokens retain strong local spatial correlations on the original grid. To exploit this structural prior, we propose the Visual Token Codec (VTC), a dual-path learned codec that separates global and patch tokens into dedicated coding paths. Global tokens are compressed with a lightweight factorized prior, whereas patch tokens are encoded on the patch-token grid using a spatial-channel context entropy model. To support intermediate-layer compression and practical rate adaptation, VTC further incorporates feature-matching supervision after subsequent ViT blocks and variable-rate modules within a single codec. Experiments on DINOv2 and SAM3 show that VTC consistently outperforms representative ViT feature coding baselines on classification, segmentation, and detection tasks. At 90% of uncompressed-feature performance, VTC reduces bitrate by 15.7x-37.4x across these tasks. We further provide intermediate-layer rate-utility analyses for practical transmission- and storage-oriented deployment scenarios.
Comments13 pages, 9 figures