arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.04778cs.AI

面向移动边缘智能体AI的扩散语言模型:基础、应用与挑战

Diffusion Language Models for Mobile Edge Agentic AI: Foundations, Applications, and Challenges

Chenqi Li, Minghui Min, Dusit Niyato, Wei Ni

首次发表
浏览论文内容

中文总结 AI 辅助

本综述研究面向移动边缘智能体AI的DLMs,分析其在边缘场景的适用性,涵盖相关技术、应用及开放问题,旨在连接DLMs特性与未来移动边缘智能的系统级需求。

中文摘要 AI 辅助

扩散语言模型(Diffusion Language Models, DLMs)为移动边缘智能体人工智能(AI)提供了一种非自回归替代方案,其通过迭代去噪而非从左到右解码来优化令牌。与基于自回归Transformer的大型语言模型(LLMs)相比,DLMs可并行更新多个不确定令牌,并在整个生成过程中利用双向上下文,从而实现超越固定顺序解码的更灵活的质量-延迟权衡。这些特性对边缘智能体极具吸引力,因为部分优化、提前退出以及约束引导的校正可降低响应延迟与通信开销,同时提升在嘈杂、不完整或动态上下文下的鲁棒性。本综述回顾了DLMs的基础,分析其在延迟、内存、能耗、带宽、隐私与可靠性约束下对边缘场景的适用性,涵盖资源高效架构、训练与推理加速、压缩、边缘/云端部署、通信感知服务、物联网(IoT)/无线应用及基于DLM的智能体评估,还探讨了长上下文状态管理、拆分推理、可信执行、多模态 grounding 及可复现基准测试等开放问题,旨在将DLMs的建模特性(包括双向性、并行优化、可控性及质量-延迟弹性)与未来移动边缘智能的系统级需求相连接。

英文摘要

Diffusion language models (DLMs) offer a non-autoregressive alternative for mobile edge agentic artificial intelligence (AI) by refining tokens through iterative denoising rather than left-to-right decoding. Compared with autoregressive Transformer-based large language models (LLMs), DLMs can update multiple uncertain tokens in parallel and exploit bidirectional context throughout the generation process, enabling more flexible quality-latency trade-offs beyond fixed sequential decoding. These properties are particularly attractive for edge agents, where partial refinement, early exit, and constraint-guided correction can reduce response delay and communication overhead while improving robustness under noisy, incomplete, or dynamic contexts. This survey reviews DLM foundations and analyzes their suitability for edge settings under latency, memory, energy, bandwidth, privacy, and reliability constraints. We cover resource-efficient architectures, training and inference acceleration, compression, edge/cloud deployment, communication-aware serving, Internet of Things (IoT)/wireless applications, and evaluation of DLM-based agents. We further discuss open issues in long-context state management, split inference, trustworthy execution, multimodal grounding, and reproducible benchmarking. The goal is to connect DLM modeling properties, including bidirectionality, parallel refinement, controllability, and quality-latency elasticity, with system-level requirements of future mobile edge intelligence.

发表机构

  • China University of Mining and Technology(中国矿业大学)
  • Nanyang Technological University(南洋理工大学)
  • Edith Cowan University(埃迪斯科文大学)

机构由 AI 辅助整理,请以论文原文为准。

↑