发表机构
University of Minho; Universidade do Porto; Laboratory for Artificial Intelligence and Computer Science (LIACC); Fraunhofer Portugal AICOS; INESCTEC; University of Coimbra(米尼奥大学; 波尔图大学; 人工智能与计算机科学实验室(LIACC); 弗劳恩霍夫葡萄牙AICOS机构; INESCTEC; 科英布拉大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对隐私敏感时间序列预测场景,通过TSTR协议构建基准测试,提出Grasynd-P方法,揭示不同生成方法的预测与隐私权衡,为隐私感知合成时间序列生成方法提供评估参考。
AI 中文摘要
隐私敏感领域的时间序列预测通常需要基于发布的数据而非原始观测值训练模型。合成时间序列生成方法主要用于数据增强,生成的序列用于补充原始训练集。这些方法在完全替代原始数据时的表现,以及发布序列带来的隐私风险,仍未得到充分探索。我们通过基准测试解决这一缺口,在“基于合成数据训练、真实数据测试(TSTR)”协议下评估合成生成方法和基于噪声的匿名化基线。我们在七个数据集上联合评估预测性能和基于距离的经验隐私风险,刻画这些目标之间的权衡关系。我们还引入Grasynd-P,这是基于图的生成器Grasynd的隐私导向扩展,整合了矩阵集成和核密度估计。我们的结果表明:(1)没有任何生成方法能完全替代原始训练数据;(2)基于噪声的匿名化实现最强隐私,但预测性能最差;(3)在该场景下,基于简单变换的生成器的预测表现优于深度生成模型;(4)Grasynd-P位于帕累托前沿,在实现有竞争力预测的同时,相比其他生成器具备更强的隐私区分度。该基准为评估和开发新的隐私感知合成时间序列生成方法建立了参考标准。
英文摘要
Time series forecasting in privacy-sensitive domains often requires training models on released data rather than original observations. Synthetic time series generation has been developed primarily for data augmentation, where generated series supplement the original training set. How well these methods perform when fully replacing the original data - and how much privacy risk the released series carry - remains underexplored. We address this gap through a benchmark evaluating synthetic generation methods and noise-based anonymization baselines under a Train on Synthetic, Test on Real (TSTR) protocol. We jointly assess forecasting performance and distance-based empirical privacy risk across seven datasets, characterizing the trade-off between these objectives. We also introduce Grasynda-P, a privacy-motivated extension of the graph-based generator Grasynda, incorporating matrix ensembling and kernel density estimation. Our results show that: (1) no generation method fully substitutes for original training data; (2) noise-based anonymization yields the strongest privacy but the worst forecasting performance; (3) simple transformation-based generators outperform deep generative models for forecasting in this setting; and (4) Grasynda-P lies on the Pareto frontier, achieving competitive forecasting with stronger privacy separation than other generators. This benchmark establishes a reference point for evaluating and developing new privacy-aware synthetic time series generation methods.