域边界:面向AI就绪科学数据的静默故障检测器
Domain Bounds as a Silent-Fault Detector for AI-Ready Scientific Data
- Oak Ridge National Laboratory(橡树岭国家实验室)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对AI科学工作流中数据复用导致的静默故障,提出将数据集上下文编码为域边界并关联数据,通过标准容器技术检测不兼容复用,在注入故障测试中捕获18例中的17例。
AI中文摘要:
我们针对一种新兴的数据型故障,其根源在于AI增强的计算科学工作流中数据复用量的不断增加。数据集在特定条件下创建;其类别、度量指标和标签依赖于该上下文,但下游用户通常无法获取该上下文,且当数据集用于训练模型时,该上下文被完全弃用。我们提出将此上下文(称为域边界)编码并关联到其所适用的数据上。通过规定有效复用的条件,域边界可标记表面有效的数据,防止其在不兼容的上下文中被复用,即使每个值均落在预期范围内。我们展示了计算科学和生物医学中的案例,其中缺失域边界会导致静默复用错误,并说明为何描述性和溯源方法无法检测这些错误。我们的检测器采用标准数据容器技术,能够捕获从已发表案例分类中提取的故障,无论数据集规模大小。针对注入的故障,它捕获了谱系和描述性记录所接受的18个域边界复用中的17个。
英文摘要:
We address an incipient type of data-based faults, driven by increasing amounts of data reuse in AI-enhanced computational science workflows. A dataset is created under specific conditions; its categories, measures, and labels depend on that context, but that context is usually unavailable to downstream users and is completely abandoned when datasets are used to train models. We propose the encoding and association of this context, called domain bounds, with the data to which it applies. By specifying the conditions for valid reuse, a domain bound flags apparently valid data from being reused in an incompatible context, even when every value falls within its expected range. We present cases from computational science and biomedicine where missing domain bounds allow silent reuse errors, and show why description and provenance approaches do not detect them. Our detector uses standard data container technologies, catches faults drawn from a taxonomy of published cases, irrespective of dataset size. Against injected faults it caught 17 of 18 domain-bound reuses that lineage and descriptive records accepted.