算法审计的开放脉络:为何人工智能评估在全球南方地区落后于其部署
Open Veins of Algorithmic Auditing: Why AI Assessment Lags Behind Its Deployment in the Global South
AI总结:
研究全球南方人工智能评估落后于部署的问题,通过十年审计实践发现评估稀疏及四个模式,得出五点经验教训,指出差距根源是资金问题,建议资助者、政府改进,强调评估设计需完善。
AI中文摘要:
人工智能在全球南方地区的部署速度与全球北方地区相当甚至更快,但人工智能治理却未能跟上步伐,且南方地区的差距更大。本文借鉴了在拉丁美洲、撒哈拉以南非洲和亚太地区长达十年的人工智能审计实践(包括对巴西的机器人劳拉进行的唯一一次已完全发表的已部署系统的第二方审计;两项已完成但未发布的国家审计,涉及儿童福利风险模型和公共就业匹配算法;十三项负责任人工智能评估;以及一次区域格局分析),记录了全球南方人工智能评估这一稀疏领域的模式,并解释了其为何如此罕见。我们发现过去十年中该地区已部署系统的第二方和第三方审计发表数量不到二十篇,而有记录的公共部门算法有数百个,国家人工智能投资达数十亿美元。我们发现了四个贯穿各领域的模式:用可预测性替代有效性的代理目标;在流行度分析下失效的性能声明;模型对从未在训练中见过的人群进行评分;以及即使去除受保护属性后仍持续存在的结构性偏差。然后,我们从北方地区的监管、语言和数据条件之外得出了关于评估、监管和资金的五点经验教训。我们认为,根本而言,差距不是能力问题,而是资金问题:能力随资金需求而动,目前没有任何行为体被要求或获得资金来问责已部署系统。最有能力弥合这一差距的行为体是该地区大多数重大人工智能背后的少数发展和慈善资助者,其资金条件可以要求进行独立评估,而目前监管机构尚未这样做。最后,我们为资助者、政府以及评估本身的设计提出了建议。
英文摘要:
Artificial intelligence is being deployed across the Global South at a pace matching or exceeding the Global North, yet AI governance has not kept pace, and the gap is far wider in the South. Drawing on a decade of AI audit practice across Latin America, Sub-Saharan Africa, and Asia Pacific (the only fully published second-party audit of a deployed system in the region, Robot Laura in Brazil; two completed but unreleased national audits, of a child-welfare risk model and a public-employment matching algorithm; thirteen Responsible AI Assessments; and a regional landscape analysis), this paper documents patterns in the sparse field of Global South AI evaluation and explains why it so rarely occurs. We count fewer than twenty published second- and third-party audits of deployed systems across the region over the past decade, against hundreds of documented public-sector algorithms and multibillion-dollar national AI investments. We find four cross-cutting patterns: proxy targets that substitute predictability for validity, performance claims that collapse under prevalence analysis, populations scored by models that never saw them in training, and structural bias that persists even after protected attributes are removed. We then draw five lessons for evaluation, regulation, and funding that emerge outside the regulatory, linguistic, and data conditions of the North. We argue the gap is not, at root, a capacity problem but a funding problem: capacity follows funded demand, and no actor is currently required, or funded, to hold deployed systems to account. The actor best placed to close it is the small number of development and philanthropic funders behind most consequential AI in the region, whose funding conditions can require independent evaluation where no regulator yet does. We close with recommendations for funders, governments, and the design of evaluation itself.