arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.17143cs.DC

EdgeCoInfer:用于设备端多模态大模型的分层协作推理

EdgeCoInfer: Hierarchical Collaborative Inference for On-Device Multimodal Large Models

Lin Tan, Songtao Guo, Mingyan Li, David K. Y. Yau

AI总结:

针对边缘环境中多模态大模型面临的挑战,提出EdgeCoInfer框架,通过模型间功能模块共享和模型内细粒度划分实现粒度自适应部署,采用混合进化分层强化学习解决问题,实验表明该框架能提高任务完成率并降低系统成本和内存消耗。

AI中文摘要:

现代移动应用主要通过执行并发多模态大语言模型(MLLMs)来提供普遍存在的智能。然而,由于多任务并发和严格耦合的硬约束,在边缘环境中满足这一需求面临重大挑战。为解决这些问题,我们提出了EdgeCoInfer框架,通过共同优化模型间功能模块共享和模型内细粒度划分来实现粒度自适应部署。我们通过混合进化分层强化学习(HE-HRL)范式解决潜在的混合整数非线性规划(MINLP)问题,该范式将用于离散模型放置的遗传算法(GA)与用于连续资源分配的软演员-评论家(SAC)智能体同步。为引导稀疏可行区域,我们引入了可行性引导的建设性执行机制,集成了建设性割步解码器与预激活剪枝以及用于稳定适应的两阶段课程策略。实验结果表明,与现有基线相比,EdgeCoInfer在高并发场景中确保了100%的任务完成率,系统成本降低了76%,内存节省了71.88%。

英文摘要:

To deliver ubiquitous intelligence, modern mobile applications increasingly execute concurrent Multimodal Large Language Models (MLLMs) on edge devices, presenting severe challenges under multi-task concurrency and tight resource constraints. To address this, we propose EdgeCoInfer, a hierarchical collaborative inference framework enabling efficient on-device MLLM inference through coarse-to-fine orchestration. Coarsely, EdgeCoInfer decomposes MLLMs into functional modules for inter-task sharing, avoiding redundant model loading. Finely, it partitions models at the neural network layer level and distributes segments across devices and servers. We jointly optimize layer partitioning, module sharing, and resource allocation under tight constraints. To tackle the non-differentiable combinatorial explosion, we propose a Hybrid Evolutionary Hierarchical Reinforcement Learning (HE-HRL) framework. HE-HRL synchronizes a gradient-free genetic algorithm for discrete partitioning and sharing decisions with a gradient-based soft actor-critic agent for continuous resource refinement. We further embed a constructive cut-step decoder with pre-act pruning and a two-phase curriculum to improve feasibility and accelerate convergence. Experimental results show that EdgeCoInfer breaks the edge memory wall and prevents catastrophic out-of-memory and task failures under high concurrency, reducing memory demand by 53.53\% and system cost by 59.86\% compared to existing methods.

↑