arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

表格数据的观测级水印与检测

Observation-Level Watermarking and Detection for Tabular Data

Dongyu Cui, Xuan Bi

arXiv 2607.10554首次发表:更新:

发表机构

University of Minnesota(明尼苏达大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对表格数据水印研究不足的问题,提出STAMP框架,可适应多种分布,开发相应检测机制并建立理论保证,经模拟研究和实际应用验证,该方法有效且鲁棒,能保持数据保真度和高检测率。

AI 中文摘要

随着生成式人工智能的发展,水印技术已广泛用于检测人工智能生成数据的真实性并保护用户和创作者的权利。虽然它在成像和文本数据等数据类型中已得到很好的应用,但表格数据的水印仍未得到充分探索。现有方法主要关注数值数据,对离散、分类和混合数据的研究较少。在这项工作中,我们提出了STAMP(单观测表格归因和标记程序),这是一种用于表格数据水印的新颖框架,可适应和保留广泛的分布。我们还开发了相应的检测机制,即使样本量小至一个也能可靠地识别水印。我们建立了渐近一致性和检测准确性的理论保证。最后,通过广泛的模拟研究和两个实际数据应用,我们证明了所提出的方法是有效且对子集化具有鲁棒性的,同时保持了数据保真度和高检测率。

英文摘要

With the development of generative AI, watermarking techniques have been widely used to detect the authenticity of AI-generated data and protect the rights of users and creators. While it is already well applied in data types including imaging and text data, watermarking tabular data is still under-explored. Existing methods primarily focus on numerical data, leaving discrete, categorical, and mixed data less studied. In this work, we propose STAMP (Single-observation Tabular Attribution and Marking Procedure), a novel framework for watermarking tabular data that can accommodate and preserve a wide range of distributions. We also develop a corresponding detection mechanism, which can reliably identify watermarks even when the sample size is as small as one. We establish theoretical guarantees for asymptotic consistency and detection accuracy. Finally, through extensive simulation studies and two real-data applications, we demonstrate that the proposed method is effective and robust to subsetting, while maintaining data fidelity and a high detection rate.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑