发表机构
Zhejiang University(浙江大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有视觉令牌通信静态策略无法适应图像内容与信道条件的问题,提出AdapToC自适应框架,协同设计令牌数量、不等保护与可靠性感知恢复,在匹配通信成本下实现4.20 dB峰值PSNR增益,达到最先进性能。
AI 中文摘要
令牌已成为多模态基础模型的统一接口,使得视觉令牌通信成为高效图像传递的自然范式。然而,现有方法通常依赖静态策略,无法同时适应图像内容和信道条件。此外,其令牌级效用目标并不一定能转化为图像重建质量的提升。在本文中,我们提出AdapToC,一种自适应的、面向重建的视觉令牌通信框架。在发送端,自适应选择器联合建模图像内容、信道状态和通信预算,以执行实例级资源分配。它不使用固定的令牌速率和保护策略,而是动态确定应传输的令牌数量,并根据令牌重要性和当前信道条件分配不同的保护级别。在接收端,自适应MaskGIT接收器将信道可靠性融入上下文令牌建模。它区分具有不同可靠性级别的令牌,保留高置信度观测,纠正可能损坏的令牌,并从接收到的证据和全局视觉上下文中迭代重建缺失内容。通过协同设计令牌数量、不等保护和可靠性感知恢复以提升图像级重建质量,AdapToC在匹配通信成本下,相比最强的静态基线实现了4.20 dB的峰值平均PSNR增益,并在所评估的视觉令牌通信方法中达到最先进的性能。
英文摘要
Tokens have become a unified interface for multimodal foundation models, making visual-token communication a natural paradigm for efficient image delivery. However, existing methods typically rely on static policies that cannot jointly adapt to image content and channel conditions. Moreover, their token-level utility objectives do not necessarily translate into improved image reconstruction quality. In this paper, we propose AdapToC, an adaptive, reconstruction-oriented visual-token communication framework. At the transmitter, an adaptive selector jointly models image content, channel state, and communication budget to perform instance-wise resource allocation. Rather than using a fixed token rate and protection policy, it dynamically determines how many tokens should be transmitted and assigns different protection levels according to token importance and current channel conditions. At the receiver, an adaptive MaskGIT receiver incorporates channel reliability into contextual token modeling. It distinguishes tokens with different reliability levels, preserves high-confidence observations, corrects potentially corrupted tokens, and iteratively reconstructs missing content from the received evidence and global visual context. By co-designing token quantity, unequal protection, and reliability-aware recovery for image-level reconstruction quality, AdapToC achieves a peak mean PSNR gain of 4.20 dB over the strongest static baseline under matched communication costs and state-of-the-art performance among the evaluated visual-token communication methods.