arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.00922cs.DC

从云端到人群:通过去中心化边缘协作实现面向检索增强生成(RAG)的大语言模型服务民主化

From Cloud to Crowd: Democratizing LLM Service with Decentralized Edge Collaboration for RAG

Jiaxing Li, Hengzhi Wang, Feng Wang, Chi Xu, Danyang Song, Ruixiao Zhang, Edith C. H. Ngai, Jiangchuan Liu

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出DEFRAG系统,通过异构边缘设备的去中心化协作优化RAG的检索与生成,缩小SLM与LLM的准确率差距,大幅降低成本并提升吞吐量,助力边缘LLM服务民主化。

中文摘要 AI 辅助

大语言模型(LLMs)的快速发展提升了对可扩展且高性价比部署方案的需求,尤其是针对移动设备和边缘设备。云端托管的LLMs功能强大,但受厂商锁定和高资源需求影响,部署成本高昂且难以扩展,导致负载下费用高、性能不稳定。近期研究聚焦于从LLMs蒸馏或剪枝得到的小型语言模型(SLMs),将其部署在资源受限的边缘设备上以降低成本、提升可扩展性。然而,基于边缘的SLMs存在知识覆盖有限的问题,且与云端LLMs相比存在明显的准确率差距。为解决这一问题,我们提出DEFRAG,这是一种面向检索增强生成(RAG)的去中心化边缘协作系统,可在异构边缘设备上优化检索与生成过程。在检索环节,DEFRAG压缩并共享知识图谱,采用混合检索方式扩展知识覆盖范围;在生成环节,DEFRAG引入优化器,针对每个查询自适应选择SLMs和RAG参数,平衡准确率与成本。我们在异构边缘测试平台上实现了DEFRAG,并在基准问答数据集上对其进行评估,还在移动路由压力、非均匀数据放置以及特定领域问答工作负载等场景下进行测试。结果表明,DEFRAG在这些更广泛的场景下能保持稳定的服务质量和成本效率,同时可缩小SLM与LLM之间的准确率差距,相比集中式服务,成本最高可降低98.4%,峰值吞吐量最高可提升97.8%。这些发现证明了DEFRAG在边缘实现面向大众的LLM服务的潜力。

英文摘要

The rapid advancement of large language models (LLMs) has increased demand for scalable and cost-effective deployment, especially for mobile and edge devices. Cloud-hosted LLMs are powerful but expensive and difficult to scale due to vendor lock-in and high resource needs, resulting in high expenses and unstable performance under load. Recent efforts focus on deploying small language models (SLMs), distilled or pruned from LLMs, on resource-constrained edge devices to reduce costs and improve scalability. However, edge-based SLMs face limited knowledge coverage and notable accuracy gap compared to cloud-based LLMs. To address this, we present DEFRAG, a decentralized edge collaboration system for retrieval-augmented generation (RAG) that optimizes both retrieval and generation across heterogeneous edge devices. For retrieval, DEFRAG compresses and shares knowledge graphs, using hybrid retrieval to expand knowledge coverage. For generation, DEFRAG introduces an optimizer that adaptively selects SLMs and RAG parameters per query, balancing accuracy and cost. We implement DEFRAG on a heterogeneous edge testbed and evaluate it on benchmark QA datasets. We also test it under mobile route stress, non-uniform data placement, and a domain-specific QA workload. The results show that DEFRAG maintains stable service quality and cost efficiency under these broader settings. Results show that DEFRAG narrows the SLM-LLM accuracy gap, while reducing cost by up to 98.4% and increasing peak throughput by up to 97.8% over centralized services. These findings demonstrate the potential of DEFRAG for democratized LLM services at the edge.

补充信息

↑