发表机构
The University of Tokyo; Trade Union University; Academy Of Finance; Ministry of Industry and Trade(东京大学; 工会大学; 财政学院; 工业与贸易部)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对合成数据用于活动决策支持时决策不一致问题,提出策略模拟保真度标准、PolicySynth框架及三轴报告标准,通过实验对比,该框架在电信和银行语料库上展现高SSF及稳定性,可靠支持决策筛选,投资回报率虽需校正但表现良好。
AI 中文摘要
决策支持系统(DSS)越来越多地对合成客户群体进行留存假设分析,因为隐私限制使得无法无限制地使用真实数据。只有当合成数据能引导管理者做出与真实数据相同的决策时,这样的系统才是可信的。然而,现行标准仅证明分布相似性,而非决策一致性,所以合成群体可能匹配所有边际分布,但仍会使营销团队做出错误决策。我们通过三项贡献弥合了这一决策一致性差距:策略模拟保真度(SSF),一种衡量合成群体与真实群体做出相同活动决策的频率的标准;PolicySynth,一个DSS框架,其生成器以生产流失评分器为条件,以对齐与决策相关的结构;以及一个由决策一致性、成员推理抗性和新记录率组成的三轴报告标准,作为最低部署质量门槛。在电信流失语料库和银行获取语料库上,PolicySynth的平均SSF分别达到0.923和0.960,种子到种子的方差比CTGAN在电信上大约紧十倍,在银行上紧2.5倍。这种稳定性是可部署的属性:在每月再训练周期之间,活动决策建议的变化最多为1.2个百分点,而CTGAN为11.5个百分点,九分之一的活动出现反向建议。一个自举基线在SSF上与PolicySynth匹配,但逐字复制真实记录且无法通过成员推理,这表明单一轴是不够的。PolicySynth可靠地支持定向活动决策筛选;其投资回报率估计与实际结果相差70%至78%,需要我们记录的数量校正。
英文摘要
Decision support systems (DSS) increasingly run retention what-if analysis on synthetic customer populations, because privacy constraints preclude unrestricted use of real data. Such a system is trustworthy only if the synthetic data lead managers to the same decisions as the real data would; yet prevailing criteria certify distributional similarity, not decision alignment, so a synthetic population can match every marginal distribution while still steering a marketing team toward the wrong campaigns. We close this decision-alignment gap with three contributions: strategy simulation fidelity (SSF), a criterion measuring how often the synthetic population yields the same go/no-go campaign decision as the real population; PolicySynth, a DSS framework whose generator is conditioned on the production churn scorer to align decision-relevant structure; and a three-axis reporting standard of decision alignment, membership-inference resistance, and novel-record rate as the minimum deployment quality gate. On a telecommunications churn corpus and a banking acquisition corpus, PolicySynth attains a mean SSF of 0.923 and 0.960, with seed-to-seed variance roughly ten times tighter than CTGAN on telecommunications and 2.5 times on banking. This stability is the deployable property: go/no-go recommendations shift by at most 1.2 percentage points between monthly retraining cycles, against 11.5 for CTGAN, a reversed recommendation on one campaign in nine. A bootstrap baseline matches PolicySynth on SSF yet copies real records verbatim and fails membership inference, evidence that no single axis suffices. PolicySynth reliably supports directional go/no-go screening; its ROI estimates diverge from real outcomes by 70 to 78% and require the volume correction we document.
Comments15 pages, 4 figures