发表机构
Sophea AI; KIEFER SA(Sophea AI; KIEFER SA)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对现代希腊语缺失于Nemotron检索模型及主流多语言基准的问题,提出Nemotron检索栈的端到端适配方案,含语料库挖掘等步骤,推出首个希腊语RAG基准HERA,适配后的模型在多项指标上显著提升。
AI 中文摘要
现代希腊语虽在法律、能源、金融和医疗等领域的检索增强生成(RAG)中具有重要应用,但并未出现在NVIDIA的Nemotron检索模型及主流多语言检索基准中。我们提出了Nemotron检索栈针对现代希腊语的端到端适配方案,涵盖语料库挖掘、合成监督、检索模型训练、重排器适配、读取器微调,以及名为HERA的新基准。研究表明,无参数的BM25基线在专业希腊语语料库上优于多款现成多语言密集检索模型;在65773个希腊语检索对上微调后,Nemotron 1B嵌入器的nDCG@10从0.362提升至0.835,大幅优于未适配版本。习得的语言能力可迁移至通用领域希腊语,不过其相较BM25的优势仍依赖领域;我们还适配了交叉编码器重排器,在各专业领域均实现持续提升。最后,我们对Nemotron 30B-A3B混合专家读取器进行LoRA微调以实现落地生成,经评估的答案正确率从29.4%提升至66.9%,同时显著提升了忠实度与引用质量。我们还推出了首个用于检索增强生成的大规模希腊语基准HERA,并发布适配后的模型与基准以支持未来希腊语RAG系统研究。
英文摘要
Modern Greek is absent from NVIDIA's Nemotron retrieval models and from major multilingual retrieval benchmarks, despite being important for retrieval-augmented generation (RAG) in legal, energy, financial, and medical applications. We present an end-to-end adaptation of the Nemotron retrieval stack for Modern Greek, including corpus mining, synthetic supervision, retrieval model training, reranker adaptation, reader fine-tuning, and a new benchmark called HERA. Our study shows that a parameter-free BM25 baseline outperforms several off-the-shelf multilingual dense retrieval models on specialist Greek corpora. After fine-tuning on 65,773 Greek retrieval pairs, a Nemotron 1B embedder improves nDCG@10 from 0.362 to 0.835 and substantially outperforms its unadapted counterpart. The learned language competence transfers to general-domain Greek, although the advantage over BM25 remains domain-dependent. We further adapt a cross-encoder reranker and demonstrate consistent improvements across specialist domains. Finally, we LoRA-tune a Nemotron 30B-A3B mixture-of-experts reader for grounded generation, increasing judged answer correctness from 29.4% to 66.9% while significantly improving faithfulness and citation quality. We also introduce HERA, the first large-scale Greek benchmark for retrieval-augmented generation, and release our adapted models and benchmark to support future research on Greek-language RAG systems.
Comments15 pages, 10 figures, 7 tables. Includes release of the HERA benchmark and Sophea Nemo RAG models