arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

印度评估:一个在印度语言和文化背景下评估大语言模型的开放基准框架

Inspect India Evals: An Open Benchmarking Framework for Evaluating Large Language Models in the Indian Linguistic and Cultural Context

Abhishek Kumar Singh, Shrey Nag, Sachita, Lipi Goel, Rajeshwar Singh Janwar

arXiv 2607.25375首次发表:更新:

AI 中文总结

Inspect India Evals是在印度语言文化背景下评估大语言模型的开放框架,基于Inspect AI平台构建,含六个基准。测试五个模型,Sarvam-M 24B和Gemma 2 27B表现最佳,各模型在多语言安全测试中表现一致,DPI安全得分有差异,框架公开可扩展。

AI 中文摘要

印度是一个拥有超过14亿人口的大国,有着数百种不同的地方传统和文化以及22种官方认可的语言。大语言模型正在印度大规模部署,但常见基准几乎都是以英语和西方为中心的,无法识别印度背景下独特的安全、公平和准确性问题。Inspect India Evals旨在填补这一空白,它基于英国AISI的Inspect AI平台构建,有六个基准。研究测试了五个参数从8B到32B的开放权重模型,Sarvam-M 24B和Gemma 2 27B表现最佳,综合印度公平指数得分80%,Sarvam-M在印度文化知识和DPI安全合规方面甚至超过了更大的32B模型。所有模型在多语言安全测试中得分为100%拒绝,而DPI安全得分从20%到100%不等。该框架是公开的,可与英国AISI注册中心配合使用,任何人都能重现或扩展这项工作。

英文摘要

India is a vast nation of over 1.4 billion people, varied by hundreds of diverse and locally specific traditions and cultures and 22 officially recognized languages. Large language models (LLMs) are now being deployed on a massive scale throughout the mainland as well as in remote villages. However, the common benchmarks - MMLU, BIG-Bench, and TruthfulQA are almost exclusively English- and Western-centric. They do not identify those safety, fairness, and accuracy failures unique to the Indian context. That is the gap Inspect India Evals seeks to fill. It is an open-source framework built on top of UK AISI's Inspect AI platform. It has six benchmarks: Multilingual MMLU across sixteen Indian languages, BharatBBQ (our adaptation of BBQ for Indian social bias), a safety evaluation for Digital Public Infrastructure, a multilingual safety test using harmful prompts in Indian languages, a multi-turn jailbreak resistance test, and an Indian cultural knowledge benchmark scored using LLM-as-judge rubrics. In this study, we tested five open-weight models ranging from 8B to 32B parameters. Sarvam-M 24B and Gemma 2 27B came out on top, both scoring 80% on the composite India Fairness Index, with Sarvam-M even beating larger 32B models on Indian cultural knowledge and DPI safety compliance. All models scored 100% refusal on Multilingual Safety, whereas DPI safety varied from 20% to 100%. The framework is public. It's built to work with the UK AISI registry. Anyone can reproduce or extend this work.

Comments19 pages, 9 figures, 7 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑