发表机构
IT Department, Agency for Financing Rural Investments (AFIR)(金融农村投资机构信息部)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对公共机构数据不能流出的限制,开发RAGAL检索增强助手。在特定语料库上验证,检索工程和微调效果好,记录单域微调问题及解决方法,还有两个发现和全嵌入器微调方法,最后介绍零流出下的替代评判方案并开源脚本。
AI 中文摘要
公共机构拥有大量敏感文件和支持工单,不能离开场所,因此完全排除了云托管语言模型。我们报告了RAGAL,这是罗马尼亚农村投资融资机构AFIR技术支持团队的检索增强助手,在三个严格限制下构建和运行:零数据流出(无外部API调用,即使是合成数据)、只读指令(助手起草,人工执行)以及仅使用一台8GB消费级笔记本电脑作为唯一的开发和训练机器。在约25000个罗马尼亚语块的语料库上(15073个已解决的支持工单和内部规范性文件),我们表明,最高效的投资是检索工程和检索器微调,而不是更大的生成器:带有意图路由的混合密集-稀疏检索将我们的内部评估从62%提高到81%,在真实工单数据上对bge-m3嵌入器进行72分钟训练后,召回率@10从0.663提高到0.850(平均倒数排名从0.489提高到0.684)。我们记录了一个常见问题:单域微调会在未触及的文档域上悄无声息地降低检索效果,低于原始基线,只有在构建每个域的评估集并使用本地生成的查询修复后才被检测到。我们报告了两个违反直觉的发现——个人身份信息屏蔽提高了生成质量,以及一种结构化的“锚点蒸馏”方案从根本上使SQL幻觉成为不可能——以及在8GB VRAM中进行全嵌入器微调的可重现方法。最后,由于零流出也排除了云评判器,我们描述了一种替代方案:在CPU上运行的744B参数模型,交互服务太慢,但在夜间批处理中可行,用作第二意见并量化了其局限性。我们为面临类似数据本地化限制的机构发布了经过清理的管道脚本。
英文摘要
Public institutions hold large volumes of sensitive documents and support tickets that cannot leave the premises, ruling out cloud-hosted language models entirely. We report on RAGAL, a retrieval-augmented assistant for the technical-support team of AFIR, the Romanian Agency for Financing Rural Investments, built and operated under three hard constraints: zero data egress (no external API calls, even for synthetic data), a read-only mandate (the assistant drafts, humans execute), and a single 8 GB consumer laptop as the only development and training machine. Over a Romanian-language corpus of ~25,000 chunks -- 15,073 resolved support tickets and internal normative documents -- we show that the highest-leverage investments were retrieval engineering and retriever fine-tuning rather than a larger generator: hybrid dense-sparse retrieval with intent routing raised our internal evaluation from 62% to 81%, and fine-tuning the bge-m3 embedder on real ticket data improved recall@10 from 0.663 to 0.850 (MRR 0.489 to 0.684) after 72 minutes of training. We document a general pitfall: single-domain fine-tuning silently degraded retrieval on the untouched document domain below the stock baseline, detected only after building a per-domain evaluation set and repaired with locally generated queries (GenQ). We report two counter-intuitive findings -- PII masking improved generation quality, and a structural "anchor distillation" scheme made SQL hallucination impossible by construction -- along with a reproducible recipe for full embedder fine-tuning in 8 GB of VRAM. Finally, since zero egress also rules out a cloud judge, we describe a substitute: a 744B-parameter model run on CPU, too slow to serve interactively but affordable in overnight batch, used as a second opinion whose limits we quantify. We release the sanitized pipeline scripts for institutions facing similar data-locality constraints.
Comments16 pages, 6 figures