arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.02949cs.AIcs.DB

关于缺失的数据层及一种潜在解决方案

On the missing data layer and a potential solution

  • SURUS

机构由 AI 辅助整理,请以论文原文为准。

Francis F Daniel, Mauro Ibañez, Francis Perelman, Marian Basti

AI总结:

针对拉丁美洲AI基础设施缺失的数据集层存在的发现与供给问题,提出以任务为核心的DataHub数据基础设施解决方案。

AI中文摘要:

拉丁美洲缺失人工智能基础设施的两个基础层:数据集层和基准层,本文聚焦于数据集层。该层面临两个相互叠加的问题:发现与供给。拉丁美洲的人工智能数据集存在,但分散在各平台且无共享索引;即便索引完善,其总量仍远低于前沿人工智能发展所需规模。本文提出DataHub:一种以任务为核心的数据基础设施,通过本体/任务?/领域?/语言?组织,具备数据集发现、元数据管理、贡献、许可及复用机制。

英文摘要:

Latin America is missing two foundational layers of AI infrastructure: the dataset layer and the benchmark layer. This paper targets the dataset layer. The dataset layer faces two compounding problems: discovery and supply. Latin American AI datasets exist but are scattered across platforms with no shared index. Even with perfect indexing, the total volume would remain far below what frontier AI development requires. We propose DataHub: a task-first data infrastructure organized through the ontology /<task?>/<domain?>/<language?>, with mechanisms for dataset discovery, metadata, contribution, licensing, and reuse.

补充信息

↑