发表机构
Concordia University; Toronto Metropolitan University(康考迪亚大学; 多伦多都会大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对合成人群生成的抽样零值与不可行结构零值问题,提出带IGP、LDR、CLAP正则化项的两阶段生成框架,提升了表格与序列出行属性合成的可行性、多样性与新颖性。
AI 中文摘要
合成人群是基于活动的出行需求模型的关键输入,但从有限调查数据生成真实人群仍具挑战性。小样本会遗漏有效属性组合(即抽样零值),生成模型还可能产生不可行的结构零值。此外,真实的合成人群必须同时捕捉静态社会人口统计属性和序列出行行为(如行程链)。本文提出一种正则化两阶段生成框架来应对这些挑战,其中正则化指额外的损失项,用于引导生成器实现更广泛的有效覆盖和更少的不可行样本。在第一阶段,带梯度惩罚的Wasserstein生成对抗网络(WGAN-GP)新增了IGP、LDR和CLAP三个正则化项,以提升表格人群合成的可行性、多样性和新颖性。在第二阶段,Transformer和LSTM-Attention模型以合成的表格属性为条件,生成序列出行属性,包括出发时间、行程目的和出行方式。我们还引入了新颖性和计数感知指标,用于评估是否恢复了有效的未见组合并以真实比例生成。结果表明,正则化模型在可行性、多样性和新颖性方面均优于普通WGAN-GP:正则化使可行性提升2.1至3.7个百分点,新颖性提升6.6至10.0个百分点,在不牺牲可行性的前提下改善了抽样零值的恢复效果,F1分数提升6.3至8.6个百分点;对于序列属性,LSTM-Attention最能匹配行程长度分布,而Transformer的整体序列F1分数更高,为90.6%,LSTM-Attention为89.1%;跨阶段验证证实,生成的出行状态与生成的行程链之间具有强一致性。
英文摘要
Synthetic populations are critical inputs for activity-based travel demand models, yet generating realistic populations from limited survey data remains challenging. Small samples miss valid attribute combinations, known as sampling zeros, and generative models may also produce infeasible structural zeros. Moreover, realistic synthetic populations must capture both static socio-demographic attributes and sequential travel behaviour, such as trip chains. This paper proposes a regularized two-stage generative framework to address these challenges, where regularization refers to additional loss terms that guide the generator toward broader valid coverage and fewer infeasible samples. In Stage 1, a Wasserstein GAN with gradient penalty is augmented with three regularization terms, IGP, LDR, and CLAP, to improve feasibility, diversity, and novelty in tabular population synthesis. In Stage 2, Transformer and LSTM-Attention models generate sequential travel attributes, including departure time, trip purpose, and travel mode, conditioned on the synthesized tabular profiles. We also introduce novelty and count-aware metrics to evaluate whether valid unseen combinations are recovered and generated in realistic proportions. Results show that regularized models outperform the vanilla WGAN-GP across feasibility, diversity, and novelty. Regularization increases feasibility by 2.1 to 3.7 percentage points and novelty by 6.6 to 10.0 percentage points, improving sampling-zero recovery without sacrificing feasibility. The F1 score improves by 6.3 to 8.6 percentage points. For sequential attributes, LSTM-Attention best matches the trip-length distribution, while Transformer achieves higher overall sequential F1, 90.6\% versus 89.1\%. Cross-stage validation confirms strong consistency between generated mobility status and generated trip chains.
Comments11 figures, 6 tables