极简RAG的模型:B1ade 335M嵌入模型与1B参数小型语言模型
Models for minimalist RAG: B1ade 335M Embedding and 1B Parameter Small Language Models
浏览论文内容
中文总结 AI 辅助
该研究提出含B1ade-embed嵌入模型与B1ade-1B小型语言模型的高效极简RAG架构,其在多QA基准及RAG评估中表现优异,还涌现出无明确监督的来源引用能力,验证了资源高效RAG的可行路径。
中文摘要 AI 辅助
RAG系统中使用的语言模型和嵌入模型通常被认为需要大规模预训练和明确的 grounding 监督。我们提出B1ade,一种高效的RAG架构,包含两个专用组件:紧凑嵌入模型和专用小型语言模型(SLM)。B1ade-embed是一个3.35亿参数的检索模型,通过5个预训练编码器的无参数融合构建,在零额外训练的情况下,在5亿参数以下的模型中取得了MTEB的最高分数;B1ade-1B是一个SLM,使用分组相对策略优化(GRPO)在低成本GPU上,基于7.23亿个标记(220万个样本)的精选上下文-问题对进行训练,其奖励仅优化答案相似性。我们的核心发现是涌现归因:尽管未接受任何明确的来源引用监督,B1ade-1B在42.4%的响应中引用了检索到的段落,比其训练分布的归因率高出5.5个百分点。这表明,在强化学习(RL)训练下,grounding 行为可以作为最大化准确性的策略涌现,无需明确的奖励工程。在标准QA基准上,B1ade-1B在PopQA上达到81.82%,在PubMedQA上达到65.8%,在FEVER上达到51.09%。在端到端RAG评估中,B1ade-1B在正确性、完整性、连贯性和忠实性方面的平均得分为0.654,比监督微调(SFT)的结果提高了10.8%,并缩小了与1.5倍于其规模的模型之间的差距。这些结果表明,战略性的模型组合和奖励设计足以实现资源高效的RAG,无需大规模预训练。
英文摘要
Language and embedding models used in RAG systems are conventionally assumed to require large-scale pretraining and explicit grounding supervision. We present B1ade, an efficient RAG architecture comprising two purpose-built components: a compact embedding model and a purpose-built SLM. B1ade-embed, a 335M parameter retrieval model constructed via parameter-free fusion of five pretrained encoders achieves top MTEB scores among sub-500M models with zero additional training, and B1ade-1B, an SLM trained on low-cost GPUs using Group Relative Policy Optimization (GRPO) on 723M tokens (2.2M examples) of curated context-question pairs with rewards that optimize only answer similarity. Our central finding is emergent attribution: despite receiving no explicit supervision for source citation, B1ade-1B cites retrieved passages in 42.4% of responses, exceeding the attribution rate of its training distribution by 5.5 percentage points. This demonstrates that grounding behavior can emerge as an accuracy-maximizing strategy under RL training, without explicit reward engineering. On standard QA benchmarks, B1ade-1B achieves 81.82% on PopQA, 65.8% on PubMedQA, and 51.09% on FEVER. In end-to-end RAG evaluation, B1ade-1B achieves an average score of 0.654 across correctness, completeness, coherence, and faithfulness, a 10.8% improvement over the SFT, while closing the gap with models 1.5x its size. These results show that strategic model composition and reward design suffice for resource-efficient RAG, without large-scale pretraining.
发表机构
- Amazon(亚马逊公司)
机构由 AI 辅助整理,请以论文原文为准。