arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

L3Cube-IndicQuest v2:用于评估大型语言模型(LLM)印度诸语言事实知识的大规模多语言基准

L3Cube-IndicQuest v2: A Large-Scale Multilingual Benchmark for Evaluating Factual Knowledge of Large Language Models Across Indic Languages

Rinit Jain, Tirthraj Mahajan, Advait Joshi, Raviraj Joshi

arXiv 2608.15535首次发表:更新:

发表机构

Pune Institute of Computer Technology; L3Cube Labs; Indian Institute of Technology Madras(浦那计算机技术学院; L3Cube实验室; 马德拉斯印度理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究构建了面向印度诸语言的大规模多语言问答基准L3Cube-IndicQuest v2,评估6种LLM的事实知识,发现前沿商业模型表现突出,开放权重模型Gemma4 31B优于Sarvam 30B。

AI 中文摘要

我们提出L3Cube-IndicQuest v2,这是一个用于评估大型语言模型(LLM)印度特定事实知识的大规模黄金标准多语言问答基准。该基准包含3471个基于课程的英语问答对,涵盖9个领域,这些问答对从教育课程、竞争性考试材料和特定领域参考书中整理而来。我们引入了一种实用的混合构建策略,将基于上下文的LLM问答生成与验证相结合,同时进行语义去重和人工验证,该策略可在保留标注质量的同时实现基准数据的规模化创建。该基准被翻译成19种印度诸语言,生成了一个公开可用的多语言数据集,包含20种语言的69420个问答对。我们在三种协议下评估了6个LLM:LLM作为评判者以及两种确定性词汇标准(精确子串匹配和单词重叠匹配)。所有三种协议产生的模型排名几乎相同,表明结果不依赖于评判者的选择。前沿商业模型领先优势显著,在所有评估的印度诸语言中,开放权重模型Gemma4 31B的表现优于针对印度诸语言优化的Sarvam 30B。

英文摘要

We present L3Cube-IndicQuest v2, a large-scale gold-standard multilingual question-answering benchmark for evaluating the India-specific factual knowledge of Large Language Models (LLMs). The benchmark comprises 3,471 curriculum-grounded English question--answer pairs spanning nine domains, curated from educational curricula, competitive examination materials, and domain-specific reference books. We introduce a practical hybrid construction strategy that combines context-grounded LLM-based question generation and validation with semantic deduplication and human verification, enabling scalable creation of benchmark data while preserving annotation quality. The benchmark is translated into 19 Indic languages, yielding a publicly released multilingual dataset of 69,420 question--answer pairs across 20 languages. We evaluate six LLMs under three protocols: LLM-as-a-judge and two deterministic lexical criteria, exact-substring and word-overlap matching. All three produce almost the same model ranking, showing that the results do not depend on the choice of judge. The frontier commercial model leads by a wide margin, and among open-weight models Gemma4 31B outperforms the Indic-specialised Sarvam 30B in every evaluated Indic language.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑