AI 中文总结
本文基于杨的压迫与结构不公理论,指出AI基准作为社会技术人工物,其现有实践会强化权力结构、延续系统性伤害,阻碍AI领域稳健且有益的发展。
AI 中文摘要
人工智能(AI)基准并非中立的评估工具,而是塑造AI领域内竞争、权力与研究优先级的社会技术人工物。基准标准化了系统评估,助力创建排行榜,以声望、引用、信任及机构影响力奖励最先进的性能。随着开发有竞争力的AI系统成本上升,这些奖励愈发集中于实力雄厚、产业资助的实验室。本文将这些关切置于艾里斯·马里昂·杨的压迫与结构不公理论框架中,认为当前基准实践可能延续系统性伤害,影响AI研究中的各类主体,契合杨提出的四种“压迫面孔”。基准文化进一步被界定为结构不公的来源,因为这些伤害源于常态化、个体可辩护的实践及网络效应,即便无明确不当行为。通过强化现有权力结构并收窄可能的研究轨迹,基准实际上可能阻碍该领域以认知上稳健且社会有益的方式推进。
英文摘要
Artificial intelligence (AI) benchmarks are not neutral tools of evaluation but socio-technical artefacts that shape competition, power, and research priorities within AI. Benchmarks standardise the assessment of systems and facilitate the creation of leaderboards that reward state-of-the-art performance with prestige, citations, trust, and institutional influence. As the costs of developing competitive AI systems rise, these rewards increasingly concentrate among powerful, industry-funded labs. This paper situates these concerns within Iris Marion Young's theories of oppression and structural injustice. It argues that current benchmarking practices may perpetuate systematic harms affecting various actors in AI research, aligning with four of Young's "faces of oppression". Benchmarking culture is further framed as a source of structural injustice, as these harms emerge from normalised, individually defensible practices and network effects, even without explicit wrongdoing. By reinforcing existing power structures and narrowing possible research trajectories, benchmarking may in fact prevent the field from advancing in epistemically robust and socially beneficial ways.
Commentsto appear in the proceedings of AIES 2026