arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

多模态大语言模型的令牌通信

Token Communication for Multimodal Large Language Model

Jingkai Ying, Zhijin Qin, Yuan Shen, Khaled B. Letaief

arXiv 2608.07279首次发表:更新:

AI 中文总结

本文针对多模态大语言模型(MLLMs)的令牌传输问题,提出适配MLLMs的令牌通信框架,通过集成神经编解码器、两阶段语义对齐训练及自适应适配器,在相同传输数据量下实现更优任务性能。

AI 中文摘要

随着Transformer架构的广泛成功,令牌(token)正成为一种新的基础信息处理单元,这一趋势在多模态大语言模型(MLLMs)中尤为明显,其中视觉和文本信息都以令牌的形式表示和处理。随着MLLMs的快速部署,令牌的高效传输变得愈发重要。本文研究在与MLLMs交互时如何减少传输数据量,同时保留其多模态理解性能。为解决该问题,我们提出了一种专为MLLMs设计的令牌通信框架。在该框架中,神经编解码器被集成到视觉令牌生成器中,以控制传输的比特数;在接收端,解码后的隐特征通过两条路径处理:解码器将图像重建为重建先验,适配器则将隐特征转换为视觉令牌并注入到视觉令牌生成器的中间层。为使注入的令牌适配MLLMs,我们进一步设计了两阶段视觉-语言语义对齐训练方案:适配器先通过蒸馏损失进行预热,再通过对齐损失与文本语义对齐;还引入了基于特征线性调制的自适应适配器,使单个适配器可支持多种编解码器速率。在多个MLLM基准上的大量模拟实验表明,在相同传输数据量下,所提方案比其他面向MLLMs的图像处理方案实现了更优的任务性能。

英文摘要

With the broad success of the Transformer architecture, token is becoming a new basic information processing unit. This trend is especially evident in multimodal large language models (MLLMs), where both visual and textual information are represented and processed as tokens. With the rapid deployment of MLLMs, the efficient transmission of tokens has become increasingly important. This paper investigates how to reduce the amount of transmitted data during interactions with MLLMs while preserving their multimodal understanding performance. To address this problem, we propose a token communication framework tailored to MLLMs. In the proposed framework, a neural codec is integrated into the vision tokenizer to control the number of transmitted bits. At the receiver, the decoded latents are processed through two paths. The decoder reconstructs image as a reconstruction prior, while the adapter converts latents into visual tokens and injects them into an intermediate layer of the vision tokenizer. To make the injected tokens suitable for MLLMs, we further design a two-stage visual-language semantic alignment training scheme. The adapter is first warmed up by a distillation loss and then aligned with textual semantics through an alignment loss. An adaptive adapter is also introduced through feature-wise linear modulation, allowing one adapter to support multiple codec rates. Extensive simulations on various MLLM benchmarks show that, under the same amount of transmitted data, the proposed scheme achieves better task performance than other image processing schemes for MLLMs.

Comments13 pages, 11 figures, 2 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑