arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.03148cs.LGcs.AI

面向移动端检索增强生成的轻量级分块选择

Lightweight Chunk Selection for Mobile Retrieval-Augmented Generation

Sicong Chang, Yidan Shen, Wen Yu, Jiefu Chen, Xin Fu, Renjie Hu

AI总结:

本文针对移动端RAG部署的成本问题,提出结合LLM侧查询意图、MoE路由信号与检索分块嵌入的轻量级分块选择方法,提升了排名1的证据选择性能。

AI中文摘要:

检索增强生成(RAG)通过整合外部知识提升大语言模型(LLM)的事实依据性,但在移动设备和边缘设备上部署RAG仍具挑战,因为检索到的上下文会增加计算量和内存占用。降低该成本的直接方式是仅保留一个检索分块用于生成,但排名最高的检索分块未必是最具证据支持性的,因为检索相似度不一定等同于证据充分性。现有的上下文缩减方法可提升上下文质量,但通常需要额外的LLM或压缩器,在严格的移动预算下成本高昂。本文将轻量级RAG分块选择作为证据对齐问题进行研究。我们的选择器结合三类互补特征源:代表LLM侧查询意图的问题隐藏状态、捕获生成器内部路由结构的混合专家(MoE)路由衍生专家信号、保留候选侧证据几何的检索分块嵌入。一个紧凑的多层感知器将这些特征映射到分块嵌入空间中的证据原型,通过余弦相似度选择与该原型最对齐的候选。为满足更严格的部署预算,我们进一步引入可选的任务感知特征选择策略以降低选择器输入维度。为支持监督评估,我们基于证据充分性而非答案字符串包含性构建语义分块正确性标签。实验表明,所提出的选择器相较于适用于移动端的基线方法,在排名1的证据选择上平均提升2.5%。这些结果表明,利用LLM侧查询表示和MoE路由信息,并将其与检索侧候选嵌入对齐,是一种适用于移动端的RAG分块选择的有效且参数高效的策略。

英文摘要:

RAG improves the factual grounding of LLM by incorporating external knowledge, but deploying RAG on mobile and edge devices remains challenging because retrieved context increases computation and memory. A direct way to reduce this cost is to retain only one retrieved chunk before generation, but the top-ranked retrieved chunk is not always the most evidence-supporting one, since retrieval similarity does not necessarily imply evidential sufficiency. Existing context-reduction methods can improve context quality, but often require additional LLMs or compressors that are costly under a strict mobile budget. In this paper, we study lightweight RAG chunk selection as an evidence-alignment problem. Our selector combines three complementary feature sources: question hidden states that represent LLM-side query intent, MoE routing-derived expert signals that capture the generator's internal routing structure, and retrieved chunk embeddings that preserve candidate-side evidence geometry. A compact multilayer perceptron maps these features to an evidence prototype in the chunk embedding space, and the candidate most aligned with this prototype is selected by cosine similarity. For stricter deployment budgets, we further introduce an optional task-aware feature selection strategy to reduce the selector input dimension. To support supervised evaluation, we construct semantic chunk-correctness labels based on evidence sufficiency rather than answer-string containment. Experiments show that the proposed selector consistently improves rank-1 evidence selection over mobile-applicable baselines by an average of 2.5%. These results suggest that using LLM-side query representations and MoE routing information and aligning them with retrieval-side candidate embedding is an effective and parameter-efficient strategy for mobile-applicable RAG chunk selection.

↑