发表机构
Wireless Systems Laboratory, Wireless Networks Research Center; National Institute of Information and Communications Technology (NICT)(无线系统实验室,无线网络研究中心; 信息通信技术研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文在软件定义无线电平台上实现了零样本语义通信,利用视觉-语言模型嵌入和余弦相似度,在减少信道使用的同时保持高分类准确率,并验证了实时可行性。
AI 中文摘要
语义通信近年来因其通过传输任务导向的表示而非原始数据来减少通信链路中传输数据量的能力而受到关注。零样本语义通信发送来自视觉-语言模型(VLM)的通用嵌入,因此同一发射器可以在无需重新训练的情况下服务于新的分类任务。然而,这一优势的大部分证据来自数值模拟。我们在软件定义无线电平台上实现了零样本语义通信:一个树莓派驱动一对模拟设备主动学习模块(ADALM)-Pluto收发器,发射器端配备图像编码器,接收器端配备文本编码器,并通过余弦相似度确定零样本分类结果。我们比较了两种VLM,CLIP和MobileCLIP,在各种信道条件下,即不同的信噪比(SNR)下进行了比较。我们验证了语义链路每幅图像使用的信道使用次数比JPEG加16进制正交幅度调制基线少9倍,并且在22.3 dB时在CIFAR-10上仍达到82%的准确率,而基线得分为0%。在交通标志识别数据集(TSRD)上,MobileCLIP在相同SNR下正确分类了98.3%的未见图像。将图像编码器卸载到神经处理单元可将编码时间缩短至每幅图像13.4毫秒,比树莓派4 CPU快49倍,使发射器处于实时预算之内。我们的实现在此https URL公开可用。
英文摘要
Semantic communication has recently gained traction for its ability to reduce the amount of data transmitted over a communication link by transmitting a task-oriented representation instead of the raw source. Zero-shot semantic communication sends a general embedding from a vision-language model (VLM), so the same transmitter can serve new classification tasks without retraining. Most evidence for this advantage, however, comes from numerical simulation. We implement zero-shot semantic communication on a software-defined radio platform: a Raspberry Pi drives a pair of Analog Devices Active Learning Module (ADALM)-Pluto transceivers, with an image encoder at the transmitter and a text encoder at the receiver, and determines the zero-shot classification results via cosine similarity. We compare two VLMs, CLIP and MobileCLIP, across various channel conditions, i.e., different signal-to-noise ratios (SNRs). We validate that the semantic link spends 9x fewer channel uses per image than a JPEG plus 16-ary quadrature amplitude modulation baseline and still reaches 82% accuracy on CIFAR-10 at 22.3 dB, where the baseline scores 0%. On the traffic sign recognition dataset (TSRD), MobileCLIP correctly classifies 98.3% of unseen images at the same SNR. Offloading the image encoder to a neural processing unit reduces encoding to 13.4 ms per image, 49x faster than a Raspberry Pi 4 CPU, placing the transmitter within a real-time budget. Our implementation is publicly available at https://github.com/thanhlexyz/zsscsdr.
Commentsaccepted at IEEE WPMC 2026