arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

不要责怪模型,验证数据:基于SMT的数据集验证评估

Don't Blame the Model, Verify the Data: An Evaluation of SMT-based Dataset Verification (Extended Version)

Sehee Park, Dominik Geißler, Andrei Aleksandrov, Kim Völlinger

arXiv 2609.20959首次发表:更新:

发表机构

Technische Universität Berlin; Fraunhofer FOKUS(柏林工业大学; 弗劳恩霍夫FOKUS研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究首次大规模实证评估基于SMT的数据集验证,发现属性类型、规范风格和编码策略显著影响求解器性能,其中提取特征列编码最优,聚合属性可扩展性提升超2000倍。

AI 中文摘要

欧盟人工智能法案要求高风险机器学习(ML)系统的数据集满足严格的质量标准,如健全性和偏差缓解。虽然可满足性模理论(SMT)求解为验证这些属性提供了一种形式化方法,但其在现实ML设置中的可扩展性仍未得到探索。为弥补这一差距,本工作首次对两个真实世界ML数据集上的基于SMT的数据集验证进行了大规模实证研究。我们系统评估了求解器性能如何受三个关键维度影响:数据质量属性的类型、规范风格和数据集编码策略。我们的研究结果表明,基于SMT的验证在实际场景中是可行的,但每个维度都会对其产生影响:属性类型设定了可处理性极限,规范风格驱动可扩展性(对于聚合属性超过2000倍),编码策略具有系统性影响,其中提取的特征列表现最佳。

英文摘要

The EU AI Act mandates that datasets for high-risk machine learning (ML) systems meet strict quality criteria such as soundness and bias mitigation. While Satisfiability Modulo Theory (SMT) solving offers a formal approach to verifying these properties, its scalability in realistic ML settings remains unexplored. To bridge this gap, this work presents the first large-scale empirical study of SMT-based dataset verification on two real-world ML datasets. We systematically evaluate how solver performance is shaped by three key dimensions: the type of data-quality property, the specification style, and the dataset encoding strategy. Our findings demonstrate that SMT-based verification is feasible for practical scenarios, but each dimension shapes it: the property type sets the tractability limit, the specification style drives scalability (exceeding $2{,}000{\times}$ for aggregate properties), and the encoding strategy has a systematic effect, with extracted feature columns performing best.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑