LLM作为低资源语言评判者:为罗马尼亚语适配Ragas与比较排序
LLM-as-a-Judge for Low-Resource Languages: Adapting Ragas and Comparative Ranking for Romanian
浏览论文内容
中文总结 AI 辅助
本文针对低资源语言罗马尼亚语,通过适配Ragas框架和引入AdminRo-Eval数据集,验证了LLM作为评判者的可行性,发现细粒度分解在忠实度上达到96%人类一致性,比较排序在答案相关性上达到90%,并确立Gemini 2.5 Pro为稳健基线。
中文摘要 AI 辅助
评估检索增强生成(RAG)系统对于低资源语言(LRLs)仍然是一个挑战,在这些语言中,标准的基于参考的指标往往表现不佳。本文通过使用下一代模型(Gemini 2.5 和 Gemini 3)适配Ragas框架,研究了“LLM作为评判者”范式在罗马尼亚语中的可行性。我们引入了AdminRo-Eval,一个由母语者标注的罗马尼亚行政文档精选数据集,作为基准测试自动化评估器的真实参考。我们比较了三种评估方法——直接评分、比较排序和细粒度分解——在忠实度、答案相关性和上下文相关性指标上的表现。我们的研究结果表明,评估策略必须针对特定指标:细粒度分解在忠实度上实现了最高的人类一致性(使用Gemini 2.5 Pro时为96%),而比较排序在答案相关性上表现更优(90%)。此外,我们证明,尽管轻量级模型在低资源语言的复杂推理中表现困难,Gemini 2.5 Pro架构为罗马尼亚语RAG评估建立了一个稳健、可迁移的基线。
英文摘要
Evaluating Retrieval-Augmented Generation (RAG) systems remains a challenge for Low-Resource Languages (LRLs), where standard reference-based metrics fall short. This paper investigates the viability of the "LLM-as-a-Judge" paradigm for Romanian by adapting the Ragas framework using next-generation models (Gemini 2.5 and Gemini 3). We introduce AdminRo-Eval, a curated dataset of Romanian administrative documents annotated by native speakers, to serve as a ground truth for benchmarking automated evaluators. We compare three evaluation methodologies - direct scoring, comparative ranking, and granular decomposition - across metrics for Faithfulness, Answer Relevance, and Context Relevance. Our findings reveal that evaluation strategies must be metric-specific: granular decomposition achieves the highest human alignment for Faithfulness (96% with Gemini 2.5 Pro), while comparative ranking outperforms in Answer Relevance (90%). Furthermore, we demonstrate that while lightweight models struggle with complex reasoning in LRLs, the Gemini 2.5 Pro architecture establishes a robust, transferable baseline for automated Romanian RAG evaluation.
发表机构
- University of Bucharest(布加勒斯特大学)
机构由 AI 辅助整理,请以论文原文为准。