AI 中文总结
该研究针对数据准备中隐私评估滞后问题,提出将其作为交互式指导问题,通过建模准备计划为确定性操作符序列,基于兼容性集语义实现推理感知隐私指导,为数据准备全程提供基础并指出构建相关系统的关键挑战。
AI 中文摘要
数据准备通常始于敏感数据,并生成可用于分析、共享或模型训练的工件。现有工作流程主要以实用性为导向,隐私通常仅在最终发布时进行评估。我们将隐私感知数据准备作为一个交互式指导问题提出。我们将准备计划建模为一系列确定性管理操作符,并询问每个步骤如何改变具有目标先验知识的观察者可用的证据。我们的语义基于兼容性集,它捕获在观察到发布的表示后对目标仍然合理的源元组。这种观点区分了消除证据的操作符和消除歧义的操作符,解释了为什么隐私效应可能是非单调的,并支持在披露预算下的前缀级反馈。结果是一个推理感知基础,用于在整个数据准备过程中指导管理者,而不是仅在最终工件产生后判断隐私。我们通过识别构建交互式、推理感知数据准备系统中的关键挑战来得出结论。
英文摘要
Data preparation often begins with sensitive data and produces a releasable artifact for analysis, sharing, or model training. Existing workflows are primarily guided by utility: a curator drops attributes, coarsens values, filters populations, and suppresses tuples until the resulting dataset appears useful and safe. Privacy, when considered, is usually evaluated only on the final release. We propose privacy-aware data preparation as an interactive guidance problem. We model a preparation plan as a sequence of deterministic curation operators and ask how each step changes the evidence available to an observer with prior knowledge about a target. Our semantics is based on compatibility sets, which capture the source tuples still plausible for the target after a released representation is observed. This view separates operators that remove evidence from those that remove ambiguity, explains why privacy effects can be non-monotone, and supports prefix-level feedback under a disclosure budget. The result is an inference-aware foundation for guiding curators throughout data preparation, rather than judging privacy only after the final artifact is produced. We conclude by identifying the key challenges in building interactive, inference-aware data preparation systems.
CommentsVLDB 2026 Workshop:15th International Workshop on Quality in Databases (QDB 2026)