arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向令牌的语义通信与预训练视觉Transformer

Token-Oriented Semantic Communication with Pretrained Vision Transformers

Jiwoong Im, Minwoo Kim, Jaeho Lee, Yo-Seb Jeon, Yongjune Kim

arXiv 2608.25410首次发表:更新:

发表机构

Pohang University of Science and Technology (POSTECH)(浦项科技大学(POSTECH))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对令牌通信的两大挑战,提出一种模块化的面向令牌的语义通信框架,利用ViT与LIC的空间对齐等技术,在ImageNet上实现了更优的速率-准确率权衡。

AI 中文摘要

令牌通信在Transformer令牌粒度上实现了语义通信原理,为资源受限边缘系统中的客户端-服务器协同推理提供了有前景的方向。然而,直接传输令牌嵌入存在两个实际挑战:通信成本高,以及跨模型特定令牌嵌入空间的互操作性有限。为解决这些挑战,我们提出了一种面向令牌的语义通信框架。在该框架中,令牌级任务相关性决定传输哪些压缩图像潜变量,从而实现无需直接传输令牌嵌入的令牌粒度传输。该框架是模块化的,协调三个预训练组件——轻量级客户端侧视觉Transformer(ViT)、学习型图像压缩(LIC)模型和大型服务器侧ViT——无需端到端训练。关键的促成因素是ViT补丁令牌与LIC潜变量之间的一对一空间对齐,这使得令牌级任务相关性可直接决定传输哪些潜变量。基于这种对齐,令牌对齐的LIC选择性传输与任务相关的潜变量,层选择性注意力展开在单次前向传播中从选定范围的注意力层估计令牌相关性,而代理令牌替换通过优化单个可学习令牌来适配冻结的服务器模型。在ImageNet上的实验表明,与近期的语义通信方案、手工编解码器和任务不可知的LIC模型相比,所提出的框架实现了更优的速率-准确率权衡。

英文摘要

Token communications realize the semantic communication principle at the granularity of transformer tokens, providing a promising direction for client--server collaborative inference in resource-constrained edge systems. However, directly transmitting token embeddings presents two practical challenges: substantial communication cost and limited interoperability across model-specific token embedding spaces. To address these challenges, we propose a \emph{token-oriented} semantic communication framework. In this framework, token-level task relevance determines which compressed image latents are transmitted, enabling token-granular transmission without directly transmitting token embeddings. The framework is modular, coordinating three pretrained components---a lightweight client-side vision transformer (ViT), a learned image compression (LIC) model, and a large server-side ViT---without end-to-end training. The key enabler is the one-to-one spatial alignment between ViT patch tokens and the LIC latent vectors, which allows token-level task relevance to directly determine which latent vectors are transmitted. Building on this alignment, token-aligned LIC selectively transmits task-relevant latents, layer-selective attention rollout estimates token relevance from a selected range of attention layers in a single forward pass, and surrogate token substitution adapts the frozen server model by optimizing a single learnable token. Experiments on ImageNet show that the proposed framework achieves a more favorable rate--accuracy trade-off than recent semantic communication schemes, hand-crafted codecs, and task-agnostic LIC models.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑