面向数据受限环境下PM预测的自适应双编码器融合的偏移感知迁移学习
Shift Aware Transfer Learning with Adaptive Dual-Encoder Fusion for PM Forecasting in Data-Limited Environments
浏览论文内容
中文总结 AI 辅助
针对数据受限环境下PM2.5预测的负迁移问题,提出偏移感知双编码器迁移框架,经实验验证其性能优于基线模型,源域知识结合目标适配可提升预测效果。
中文摘要 AI 辅助
当目标域观测数据有限且源域与目标域的统计特性存在差异时,细颗粒物(PM2.5)的短期预测仍然存在困难。在这些场景下,仅基于本地数据训练的模型可能无法捕捉复杂的时间动态,而直接迁移学习则可能导致负迁移。本研究开发了一种偏移感知双编码器迁移框架,该框架将源域知识与目标特定表示学习相结合。源编码器使用来自美国10个监测点的每小时观测数据进行预训练,随后在台湾77个站点的两年每小时观测数据上,采用按时间顺序划分的训练-验证-测试协议对该框架进行适配与评估。在四个主要基线中,冻结源编码器的双编码器模型取得了最佳性能,其均方误差(MSE)为21.8960,平均绝对误差(MAE)为3.1597,决定系数(R²)为0.8725;与TL-v1相比,MSE降低约7.1%,与TL-v2相比,MSE降低约4.1%。消融分析显示,移除台湾特定分支会导致性能出现最大幅度的下降。允许源编码器适配则产生了最佳整体结果,其MSE为21.6575,MAE为3.1383,R²为0.8739。SHAP分析表明,预测主要由近期PM2.5观测值以及与污染物传输和扩散相关的气象变量驱动。这些结果表明,当保留目标特定信息且允许迁移后的表示在目标监督下进行适配时,源域知识最为有效。
英文摘要
Short-horizon forecasting of fine particulate matter (PM2.5) remains difficult when observations from the target domain are limited and the statistical properties of the source and target domains differ. In these settings, models trained only on local data may not capture complex temporal dynamics, while direct transfer learning can result in negative transfer. This study develops a shift-aware dual-encoder transfer framework that combines source-domain knowledge with target-specific representation learning. The source encoder was pretrained using hourly observations from 10 U.S. monitoring locations. The framework was then adapted and evaluated using two years of hourly observations from 77 stations in Taiwan under a chronological train-validation-test protocol. Among the four principal baselines, the frozen-source dual-encoder model achieved the best performance, with MSE = 21.8960, MAE = 3.1597, and R^2 = 0.8725. This corresponds to an MSE reduction of approximately 7.1% relative to TL-v1 and 4.1% relative to TL-v2. The ablation analysis showed that removing the Taiwan-specific branch caused the largest decline in performance. Allowing the source encoder to adapt produced the best overall result, with MSE = 21.6575, MAE = 3.1383, and R^2 = 0.8739. SHAP analysis indicated that predictions were driven mainly by recent PM2.5 observations and meteorological variables related to pollutant transport and dispersion. These results suggest that source-domain knowledge is most effective when target-specific information is preserved and the transferred representation is allowed to adapt under target supervision.