arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.14661cs.AI

SmartRAG:用于移动设备的基于原生图的RAG

SmartRAG: Native Graph-Based RAG for Mobile Device

Zhihan Jiang, Meng Li, Shenghao Liu, Keran Li, Ruiben Zhou, Wei Wang, Xianjun Deng, Shuai Wang, Haipeng Dai

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对移动设备部署大语言模型的问题,提出SmartRAG框架,围绕四个模块组织智能助手,核心是可持续学习的EvoNER,知识存于MRGraph,通过混合管道检索,实验表明其在多跳推理性能上有竞争力且能在普通智能手机上运行。

中文摘要 AI 辅助

在移动设备上部署大语言模型作为个人助手需要隐私、低延迟和离线可用性,但巨型模型的计算成本与严格的边缘硬件预算相冲突。我们认为仅靠模型压缩无法解决这种矛盾,需要将设备上的智能分解为互补的功能角色。我们提出了SmartRAG,这是一个完全在设备上的框架,围绕四个协调模块——感知、记忆、聚焦和思考来组织智能助手。SmartRAG的核心是EvoNER,它是一个可持续学习的命名实体识别器,通过教师提炼更新逐步扩展其标签库。提取的知识存储在MRGraph中,并在查询时通过结合图遍历、词汇匹配和密集语义搜索的混合管道进行检索。仅在高价值语义操作时调用设备上的大语言模型,以限制推理成本。在四个问答基准测试上的实验表明,具有量化的17亿参数主干的SmartRAG实现了与高达18倍大的模型具有竞争力的多跳推理性能,同时完全在普通智能手机上运行,内存和延迟在实际范围内。

英文摘要

Deploying large language models (LLMs) as personal assistants on mobile devices demands privacy, low latency, and offline availability, yet the computational cost of giant models clashes with strict edge-hardware budgets. We argue that this tension cannot be resolved by model compression alone; it requires decomposing on-device intelligence into complementary functional roles. We present SmartRAG, a fully on-device framework that organizes an intelligent assistant around four coordinated modules -- Perception, Memory, Focus, and Thinking. At the core of SmartRAG is EvoNER, a continually learnable named-entity recognizer that incrementally expands its label inventory through teacher-distilled updates, enabling the system to absorb previously unseen entity types without retraining the backbone LLM. Extracted knowledge is stored in MRGraph, a three-layer provenance-preserving knowledge graph, and retrieved at query time through a hybrid pipeline combining graph traversal, lexical matching, and dense semantic search. The on-device LLM is invoked only for high-value semantic operations -- labeling, planning, and answer synthesis -- keeping inference costs bounded. Experiments on four QA benchmarks (TriviaQA, Natural Questions, HotpotQA, MultiHopQA) show that SmartRAG with a quantized 1.7B-parameter backbone achieves multi-hop reasoning performance competitive with models up to 18$\times$ larger, while running entirely on commodity smartphones within practical memory and latency envelopes.

发表机构

  • Nanjing University(南京大学)
  • HUST(华中科技大学)
  • Southeast University(东南大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑