arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.29284cs.CLcs.IR

通过针对性检索器微调实现乌兹别克斯坦法律RAG的云端与本地部署

Cloud and On-Premises Deployment of Uzbek Legal RAG via Targeted Retriever Fine-Tuning

  • Metric AI Lab

机构由 AI 辅助整理,请以论文原文为准。

Tatul Danielyan, Mariam Avetisyan, Hrant Davtyan

AI总结:

针对乌兹别克斯坦低资源语言的法律RAG部署场景,构建两个领域基准,训练开源文本嵌入器UTE-1,提炼实用指南并发布相关资源。

AI中文摘要:

将大语言模型部署用于法律问答会面临通用排行榜未覆盖的挑战,尤其针对低资源语言及严格的运营约束场景。本文报告了乌兹别克斯坦检索增强生成(RAG)法律助手的构建与运营,该助手需在两种模式下运行:一是托管云服务,需在每token成本上限内最大化回答质量;二是面向本地部署的客户端,其法律数据不得离开自身基础设施,因此我们受限于有限本地硬件上的开源权重模型及延迟约束。由于该场景缺乏评估基准,我们构建了两个领域基准:一是包含178条专家标注的法律查询及对应黄金条款跨度的检索基准;二是包含504条专家整理的问答对的端到端基准,这些问答对由大语言模型评判器打分,且我们验证了其评分与人工评判及独立家族评判器的一致性。在两种模式下应用这些基准后,我们发现开源模型与专有模型的性能差距很小,且可通过微调低成本缩小。因此,我们训练了UTE-1,这是一款针对乌兹别克斯坦的开源模型中性能领先的文本嵌入器。我们还证明,通过微调缩小性能差距既不切实际(因长上下文法律问答需大量硬件资源)也无必要(因法律条文频繁变更),QLoRA实验的负面结果支持了这一结论。我们从生产环境服务真实用户的系统中提炼出类似部署的实用指南,并发布了基准、评估代码及微调后的嵌入器UTE-1(可在指定链接获取),以支持未来低资源法律自然语言处理的研究工作。

英文摘要:

Deploying large language models for legal question answering raises challenges that general-purpose leaderboards do not capture, particularly for low-resource languages and under hard operational constraints. We report on building and operating a retrieval-augmented (RAG) legal assistant for Uzbek that must run in two regimes: a managed cloud service that maximizes answer quality within a per-token cost ceiling, and an on-premises deployment for clients whose legal data may not leave their infrastructure, restricting us to open-weight models on limited local hardware under latency constraints. Because no evaluation existed for this setting, we build two domain benchmarks: a retrieval benchmark of 178 expert-annotated legal queries with gold provision spans, and an end-to-end benchmark of 504 expert-curated question--answer pairs scored by an LLM judge whose ratings we validate against human judgments and against an independent-family judge. Applying these benchmarks under each regime, we find the open-versus-proprietary gap is small and cheaply closed by fine-tuning. Therefore, we train UTE-1, which is a state-of-the-art text embedder among open models for Uzbek. We also demonstrate that closing the performance gap via fine-tuning is both impractical due to the intensive hardware demands of long-context legal Q\&A and unnecessary, given that legal acts change frequently. We support this by reporting a negative result from a QLoRA experiment. We distill practical guidance for similar deployments, drawn from a system serving real users in production. We release our benchmarks, evaluation code and the fine-tuned embedder (UTE-1) \href{https://metric-ai-lab.github.io/Uzbek-Legal-RAG/}{at this https URL} to support future work on low-resource legal NLP.

补充信息

↑