arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

缺失的基准层及一种潜在解决方案

On the missing benchmarks layer and a potential solution

Francis F Daniel, Mauro Ibañez, Francis Perelman, Marian Basti

arXiv 2608.02996首次发表:更新:

发表机构

SURUS(苏鲁斯)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对拉丁美洲缺失本土AI开发基准层的问题,本文提出EvalsHub及首个区域实例LatamBoard,该开放基准基础设施可用于多主体的AI评估,兼具审计与优化功能。

AI 中文摘要

拉丁美洲缺少本土AI开发的基础层:基准层。基准层具备其他层无法实现的两项功能——依据区域社会需求审计AI系统,以及在经济相关环境中指导AI优化。若无该层,公共机构无法独立评估境外AI系统,企业也无法优化AI系统以达到SOTA性能来解决本地问题。缺失该层的代价是双重的:对日益成为关键基础设施的技术丧失可审计性与优化方向。本文提出EvalsHub,其首个区域实例为LatamBoard——这是一个开放的、以任务为核心的基准基础设施,高校、公共机构、专业社区及企业可在此发布、执行、对比和维护模型、工作流及智能体的评估。该基础设施一经搭建即可长期使用:新AI系统推出时机构可重新运行评估,系统每次更新后行业团队也可重新运行评估,其设计秉持开放原则,构建则由激励机制驱动。

英文摘要

Latin America is missing a foundational layer for native AI development: the benchmark layer. The benchmark layer does two things no other layer can - it audits AI systems against regional social requirements and it directs AI optimization in economically relevant environments. Without it, public institutions cannot independently evaluate foreign AI systems, and companies cannot optimize AI systems to solve local problems with SOTA performance. The cost of the missing layer is dual: a loss of auditability and a loss of optimization direction over a technology that is increasingly critical infrastructure. We propose an EvalsHub, with LatamBoard as its first regional instance - an open, task-first benchmark infrastructure where universities, public institutions, professional communities, and companies can publish, execute, compare, and maintain evaluations across models, workflows, and agents. Built once, measured forever - re-run by institutions as new AI systems ship and by industry teams after every system change. Open by design and incentive-driven by construction.

Comments9 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑