发表机构
UNIST; HUFS; JP Morgan AI Research(蔚山国立科学技术院; 韩国外国语大学; 摩根大通人工智能研究部)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究合成序列表格数据生成模型的时间保真度问题,提出分类法引导评估协议,通过数据集的四个属性确定评估维度,测量多项指标,应用于多数据集多模型,发现传统与时间评估排名差异大,强调应在时间轴测量时间保真度。
AI 中文摘要
合成序列表格数据越来越多地用于隐私保护数据共享,但生成器可能会在发出向后运行或重复的时间戳时,以及在沿真实实体未遵循的路径发送实体时,再现每一个边缘和每一个外键关系。传统的表格评估将记录汇总到静态分布中,对这些问题视而不见。我们提出了一种用于时间保真度的分类法引导评估协议,其中适用的测量由数据决定而非预先固定。每个数据集首先根据四个属性进行表征:时间如何表示、观测是否定期采样、轨迹是否相互依赖以及模式如何将实体与其历史联系起来。这些属性决定了哪些评估维度是有意义的。该协议随后测量时间戳有效性、对齐时间点的横截面结构、实体内动态以及随时间变化的关系结构,并在轨迹而非孤立行上重新进行效用和隐私评估。我们将该协议应用于跨越六个领域的13个数据集上的八个生成模型。传统评估下的排名与时间评估下的排名有很大差异,且产生的失败是架构相关而非随机的。因此,必须在时间轴本身上测量时间保真度,而不是从汇总记录分布中推断。
英文摘要
Synthetic sequential tabular data are increasingly used for privacy-preserving data sharing and research, yet conventional tabular metrics often overlook temporal structure. Existing single-table and relational evaluation protocols largely collapse records into static distributions, leaving key temporal properties insufficiently evaluated. We introduce Seq2Synth, a unified benchmark for assessing these properties. Its taxonomy characterizes temporal and schema properties to determine applicable evaluations, covering timestamp, cross-sectional, longitudinal, and structural fidelity, alongside trajectory-aware utility and privacy. Across seven core datasets from a 13-dataset benchmark and eight generators, models with near-perfect static fidelity still violate basic temporal constraints, producing duplicate timestamps, irregular intervals, and incomplete observation grids. Moreover, static and temporal-aware rankings diverge substantially, showing that temporal fidelity must be evaluated directly rather than inferred from static or relational scores. Project page and online appendices are available at: https://seq2synth.github.io/.
Comments26 pages, 10 figures, 25 tables. Extended version of the paper accepted at CIKM 2026. The conference proceedings version contains Appendix A only