arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.18117cs.AIecon.GNq-fin.EC

立场:AI排行榜对全球南方的服务不足——来自印度的案例研究

Position: AI Leaderboards Are Underserving the Global South: A Case Study from India

Sourav Banerjee, Saikat Saha

首次发表
浏览论文内容

中文总结 AI 辅助

该文以印度为案例,指出AI排行榜因缺乏独立治理等制度设计缺陷,未纳入全球南方高质量区域基准,建议建立带独立治理的区域排行榜以服务全球南方。

中文摘要 AI 辅助

本立场论文指出,AI排行榜在结构上不适合服务全球南方,因为它们缺乏独立治理、利益冲突政策以及指标演进机制。障碍并非缺少数据,高质量的区域基准已存在:印度的IndicSUPERB、MILU和LAHAJA,非洲的IrokoBench,阿拉伯语的AlGhafa,障碍在于制度设计。全球排行榜未纳入这些基准,也无治理机制强制要求其纳入。当全球北方的付费客户受影响时,商业压力会纠正排行榜的缺陷,而全球南方缺乏同等影响力。若无治理,影响印地语、斯瓦希里语或阿拉伯语使用者的缺陷会作为已记录但未解决的缺口长期存在。以印度(14亿人口、22种法定语言、拥有高质量基准但缺乏可信聚合)为案例,我们报告了对58名AI从业者咨询的结果,他们一致倾向于正式治理和基于披露的冲突管理。解决方案并非增加数据,而是建立更好的制度:从一开始就具备独立治理的区域排行榜。

英文摘要

This position paper argues that AI leaderboards are structurally ill-suited to serving the Global South because they lack independent governance, conflict-of-interest policies, and mechanisms for metric evolution. The barrier is not missing data; high-quality regional benchmarks already exist: IndicSUPERB, MILU, and LAHAJA for India; IrokoBench for Africa; AlGhafa for Arabic. The barrier is institutional design. Global leaderboards do not include these benchmarks, and no governance mechanism compels them to do so. Commercial pressure corrects leaderboard failures when paying customers in the Global North are affected. The Global South lacks equivalent leverage. Without governance, failures affecting Hindi, Swahili, or Arabic speakers persist indefinitely as documented but unaddressed gaps. Using India as a case study (1.4 billion people, 22 scheduled languages, high-quality benchmarks, but no trusted aggregation), we report findings from a consultation with 58 AI practitioners showing consistent preference for formal governance and disclosure-based conflict management. The solution is not more data but better institutions: regional leaderboards with independent governance from the start.

补充信息

↑