arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.00368eess.SP

面向智能体原生的任务导向通信:联合令牌压缩编码与调制

Agent-Native Task-Oriented Communication with Joint Token Compression Coding and Modulation

Zhuoran Xiao, Yihang Huang, Tianyu Jiao, Xiaohua Xu, Yin Xu

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对无线令牌通信系统的原生设计缺口,提出AI原生语义通信框架JTCM,通过两阶段训练实现任务导向的令牌传输,可降低传输开销并提升任务精度与鲁棒性。

中文摘要 AI 辅助

随着大型基础模型使智能体在各行业普及并成为智能系统的核心角色,后香农时代对通信范式进行根本性反思,转向AI原生、以智能体为中心的设计已不可避免。其中一个关键转变是,大型语言模型(LLM)原生处理的最小语义单元令牌(token)应取代比特成为通信的基本单元。然而,现有LLM领域的研究假设高速有线链路上的令牌传输是无损的,很大程度上忽略了无线环境固有的空中接口开销和信道失真,缺乏无线令牌通信系统的原生设计。为弥合这一差距,我们提出一种创新的令牌收发架构设计,以支持任务导向的令牌传输。具体而言,我们提出JTCM(Joint Token Coding and Modulation),这是一种AI原生的语义通信框架,联合优化令牌表示、信道编码与调制,直接最大化下游任务性能。相应地,我们提出两阶段训练方案:第一阶段预训练令牌编解码器对,实现语义保留的压缩与重建;第二阶段在特定下游任务下与多模态基础模型端到端微调,达成任务感知优化。大量实验表明,在带宽和信噪比(SNR)受限的无线信道中,与最先进的基线方法相比,JTCM显著降低了传输开销,同时提升了任务精度与鲁棒性。

英文摘要

As large foundation models empower agents to become pervasive across industries and emerge as central actors in intelligent systems, a fundamental rethinking of communication paradigms toward AI-native, agent-centric designs in the post-Shannon era becomes inevitable. One essential shift is that tokens, which are the minimal semantic units natively processed by large language models (LLMs), should replace bits as the fundamental unit of communication. However, existing works in the LLMs field assume lossless token transmission over high-speed wired links and largely neglect the air-interface overhead and channel distortions inherent in wireless environments, lacking a native design for wireless token communication systems. To bridge this gap, we propose an innovative design for a token transmitter-receiver architecture that facilitates task-oriented token transmission. Specifically, we propose JTCM (Joint Token Coding and Modulation), an AI-native semantic communication framework that jointly optimizes token representation, channel coding, and modulation to maximize downstream task performance directly. Correspondingly, we propose a two-stage training scheme. In the first stage, the token encoder-decoder pair is pre-trained to enable semantic-preserving compression and reconstruction. In the second stage, it is fine-tuned end-to-end with a multi-modal foundation model under specific downstream tasks to achieve task-aware optimization. Extensive experiments demonstrate that JTCM significantly reduces transmission overhead while enhancing task accuracy and robustness compared to state-of-the-art baselines in bandwidth- and SNR-constrained wireless channels.

补充信息

↑