基于基准的公开基准印度基础模型对比评估:能力与评估成熟度框架
Benchmark-Based Comparative Assessment of Publicly Benchmarked Indian Foundation Models: A Capability and Evaluation-Maturity Framework
浏览论文内容
中文总结 AI 辅助
本文提出能力与评估成熟度框架,对公开基准的印度基础模型与全球同类模型对比评估,发现其在传统基准表现尚可但参与新评估不足,还提出基准成熟度指数,指出能力差距或源于评估生态问题。
中文摘要 AI 辅助
政府日益加大对本土基础模型的资金投入,以增强国家AI能力、数字主权和多语言计算能力。然而,基准报告不一致、专有评估方法以及模型快速迭代的问题,给评估这类国家生态系统的进展带来了困难。本文针对公开基准的印度基础模型,与全球前沿及同等规模模型开展结构化的、基于基准的对比评估,覆盖八大能力领域:通用推理、编码与软件工程、智能体AI与计算机使用、网络安全、视觉与图像理解、视频与多模态理解、科学研究以及印度语言能力。仅使用公开报告的基准结果,研究发现印度模型在MMLU、MATH-500等成熟基准上取得了不错的分数,但这些基准如今已被广泛认为饱和,前沿模型开发者不再报告相关结果;印度模型参与较新的智能体及领域专项评估的频率低得多,且不同印度机构的基准参与度差异极大,在受调查模型中,Sarvam AI的基准覆盖范围大幅领先。研究提出了探索性的四维基准成熟度指数(BMI),从标准化、参与度、独立验证和国家覆盖度四个维度对每个能力领域打分,结果显示BMI能细化甚至修正纯描述性审查得出的成熟度判断。研究认为,现有公开记录中的诸多明显能力差距,无法与评估生态系统差距区分,这对国家AI项目设计监测和资助标准具有直接意义。
英文摘要
Purpose: Governments increasingly fund indigenous foundation models to strengthen national AI capability, digital sovereignty, and multilingual computing. This paper assesses India's foundation-model ecosystem and examines whether apparent capability gaps in public benchmark evidence may also reflect gaps in evaluation maturity. Approach: The paper presents a structured, benchmark-based comparative assessment of Indian foundation models against global frontier and comparable-scale models across eight capability domains: general-purpose reasoning, coding and software engineering, agentic AI and computer use, cybersecurity, vision and image understanding, video and multimodal understanding, scientific research, and Indic language capability. Using only publicly reported results, it proposes an exploratory four-dimension Benchmark Maturity Index (BMI), scoring each domain on standardization, participation, independent verification, and national Findings: Indian models achieve strong scores on established benchmarks such as MMLU and MATH-500. However, these are now widely regarded as saturated, and frontier developers no longer report them. Indian models participate far less frequently in newer, agentic, and domain-specialized evaluations, and participation is highly uneven across organizations. Sarvam AI reports the broadest coverage by a substantial margin. The BMI refines, and in some cases revises, the maturity judgments a purely descriptive review would produce. Practical implications: Many apparent capability gaps cannot be distinguished, on available evidence, from evaluation-ecosystem gaps, with direct implications for how national AI programs should design monitoring and funding criteria. Originality: The paper proposes BMI as a reusable instrument for scoring evaluation-ecosystem maturity at the domain level and demonstrates its application to the Indian foundation-model ecosystem.
发表机构
- Unique Identification Authority of India(印度唯一标识管理局)
机构由 AI 辅助整理,请以论文原文为准。