arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.12611stat.CO

tidysynthesis:用于合成数据生成的元包

tidysynthesis: a Meta-Package for Synthetic Data Generation

Aaron R. Williams, Jeremy Seeman, Gabriel Morrison

首次发表
浏览论文内容

中文总结 AI 辅助

研究旨在解决合成数据生成时设计选择难的问题,核心方法是引入tidysynthesis元包,其贡献在于实现建模框架与隐私方法更好互操作性,提供通用语法方便用户指定和迭代算法,还展示了功能、扩展性并给出端到端示例。

中文摘要 AI 辅助

合成数据生成使数据管理者能更轻松地共享数据集,限制对机密数据集中数据主体进行泄露性推断的可能性。生成合成数据需做出众多设计选择,但多数现有开源软件未能提供有效进行此类设计选择的通用软件基础设施。本文介绍了tidysynthesis,这是一个用于合成数据生成的元包,能实现现有建模框架与统计数据隐私方法之间更好的互操作性。它通过提供通用语法让用户更灵活地指定和迭代合成数据算法,轻松创建和修改合成数据生成管道。我们展示了tidysynthesis的功能和可扩展性,并提供了使用美国社区调查数据进行合成数据生成的端到端示例。

英文摘要

Synthetic data generation enables data curators to more easily share datasets that limits the potential for disclosive inferences about data subjects in confidential datasets. Generating synthetic data requires navigating numerous design choices; however, most existing open source software fails to provide common software infrastructure for making such design choices efficiently. In this paper, we introduce tidysynthesis, a meta-package for synthetic data generation that enables better interoperability between existing modeling frameworks and statistical data privacy methods. tidysynthesis allows users more flexibility to specify and iterate on synthetic data algorithms by providing a common syntax to easily create and modify synthetic data generation pipelines. We demonstrate the features and extensibility of tidysynthesis, as well as provide end-to-end examples for synthetic data generation using data from the American Community Survey

补充信息

↑