arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.11058cs.LGcs.DC

EMMI:通过融合表示压缩实现通信高效的多模态大语言模型推理的边缘多模态智能

EMMI: Edge Multi-Modal Intelligence for Communication-Efficient MLLM Inference via Fused Representation Compression

Motahare Mounesan, Irfan Khan

首次发表
浏览论文内容

中文总结 AI 辅助

EMMI通过边缘端模态编码、融合与压缩,传输紧凑表示而非原始数据,实现通信高效的MLLM推理,减少32倍通信量并降低3.4倍延迟。

中文摘要 AI 辅助

多模态大语言模型(MLLMs)的最新进展通过支持跨异构传感器模态(如视觉、文本和遥测数据)的推理,为边缘智能开辟了新的机遇。然而,由于现代MLLMs对计算、内存和通信的巨大需求,在资源受限的边缘平台上部署这些能力仍然具有挑战性。不同于传输原始传感器观测或在中间层对神经网络进行划分,边缘多模态智能(EMMI)在边缘设备和服务器资源之间传输紧凑表示,从而实现通信高效的边缘MLLM推理。为实现这一目标,EMMI在边缘执行模态特定编码、跨模态表示融合和学习压缩,仅将紧凑的潜在表示传输至服务器端资源以进行高容量MLLM推理。这种以表示为中心的设计降低了通信开销,保护了本地数据隐私,并在异构边缘设备和服务器端MLLM之间提供了固定大小的接口。在代表性多模态基准上的评估表明,EMMI可将通信负载减少32倍,同时保持相当的下游准确率,在带宽受限的边缘条件下,估计端到端推理延迟最多可降低3.4倍。

英文摘要

Recent advances in multimodal large language mod- els (MLLMs) have opened new opportunities for edge intelligence by enabling reasoning across heterogeneous sensor modalities, such as vision, text, and telemetry data. However, deploying these capabilities on resource-constrained edge platforms remains challenging due to the substantial computational, memory, and communication demands of modern MLLMs. Rather than transmitting raw sensor observations or partitioning neural networks at intermediate layers, Edge Multi-Modal Intelligence (EMMI) communicates a compact representation between edge devices and server resources, enabling communication-efficient edge MLLM inference. To achieve this, EMMI performs modality-specific encoding, cross-modal representation fusion, and learned compression at the edge, transmitting only a compact latent representation to server-side resources for high-capacity MLLM reasoning. This representation-centric design reduces communication overhead, preserves local data privacy, and provides a fixed-size interface between heterogeneous edge devices and server-side MLLMs. Evaluation on a representative multimodal benchmark demonstrates that EMMI can reduce the communication payload by 32x while maintaining comparable downstream accuracy, resulting in up to a 3.4x reduction in estimated end-to-end inference latency under bandwidth-constrained edge conditions.

发表机构

  • Texas A&M University(德克萨斯A&M大学)

机构由 AI 辅助整理,请以论文原文为准。

↑