发表机构
Portland State University(波特兰州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出一种机器学习框架,利用GIS衍生的建成环境特征预测交叉口行人流量,通过特征选择和梯度提升显著降低预测误差,优于传统GLM基线。
AI 中文摘要
交通部门需要整个道路网络的行人流量估计,以优先安排安全投资,但人工计数成本高昂,且仅覆盖一小部分交叉口。我们提出了一种机器学习流程,利用来自开放GIS数据的建成环境、土地利用和街道网络特征,预测俄勒冈州波特兰市101个城市交叉口的2小时下午高峰行人流量。从实践中使用的负二项GLM出发,我们增加了特征选择、计数感知的梯度提升和重复交叉验证,通过四种交叉验证策略下RMSE、MAPE和SMAPE的综合排名选择一个配置。获胜者是一种基于直方图的梯度提升模型,采用泊松损失和L1 Lasso特征选择,将交叉验证的RMSE比GLM基线降低了12%(从89.8降至78.7),并将留出集RMSE降低了19%(从108.0降至87.9)。代码已在GitHub上发布。
英文摘要
Transportation agencies need pedestrian volume estimates across entire road networks to prioritize safety investments, yet manual counts are expensive and cover only a small share of intersections. We present a machine learning pipeline that predicts 2-hour PM peak pedestrian volume at 101 urban intersections in Portland, Oregon, from built-environment, land-use, and street-network features drawn from open GIS data. Starting from the Negative Binomial GLM used in practice, we add feature selection, count-aware gradient boosting, and repeated cross-validation, selecting one configuration by a combined rank over RMSE, MAPE, and SMAPE across four cross-validation strategies. The winner, a histogram-based gradient boosting model with Poisson loss and L1 Lasso feature selection, reduces cross-validated RMSE by 12% over the GLM baseline (89.8 to 78.7) and holdout RMSE by 19% (108.0 to 87.9). Code is released on GitHub.