arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

立场:每个真实数据都是人为构建,而非客观真理

Position: Every Ground Truth is a Human Construction, not an Objective Truth

Charlotte Högberg, Ericka Johnson, Kiri L. Wagstaff

arXiv 2607.09668首次发表:更新:

AI 中文总结

探讨机器学习中真实数据集非客观真理而是人为构建,主张提高‘情境可靠性’,关注其构建能提升透明度、问责制等,利于阐明模型优缺点及更好运用数据集与模型。

AI 中文摘要

真实数据集在机器学习模型的训练和评估中作为参考值发挥着基础性作用。本立场文件认为,真实数据并非自然给定的中立客观度量,而是由人类和技术安排构建而成。我们认为,机器学习社区将受益于阐明和讨论这些往往无形或未报告的选择,并承认参考数据集是偶然的,而非普遍的。关注真实数据的情境依赖性本质,通过对数据集及其塑造的模型在何处、何时以及如何能最佳使用形成更明智的观点,可提高可靠性。我们主张提高‘情境可靠性’,包括阐明模型的局限性和优势及其真实性主张。最后,更多关注真实数据的构建可支持透明度、问责制和跨学科工作。

英文摘要

Ground truth datasets play a fundamental role as reference values in the training and evaluation of machine learning models. This position paper argues that ground truths are not neutral objective measurements that are naturally given, but instead that they are constructed by arrangements of humans and technologies. We argue that the ML community will benefit from articulating and discussing these often invisible or unreported choices and acknowledging that reference data sets are contingent, not universal. Focusing on the situated and context-dependent nature of ground truths can improve reliability by enabling a better informed perspective on where, when, and how the datasets, and the models they have shaped, can best be used. We argue for increasing `situated reliability' which includes articulating the limits and strengths of models and their truth claims. Finally, paying more attention to the construction of ground truths can support transparency, accountability, and interdisciplinary work.

Comments13 pages, 1 figure. To be published in Proceedings of the 43 rd International Conference on Machine Learning, Seoul, South Korea. PMLR 306, 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑