arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

住宅能源估算的全特征与有限输入机器学习:现实输入约束下RECS与ResStock的对比分析

Full-Feature versus Limited-Input Machine Learning for Residential Energy Estimation: A Comparative Analysis of RECS and ResStock Under Realistic Input Constraints

Aditya Ramnarayan, Fatih Evren, Patti Gunderson, Samuel Rosenberg

arXiv 2608.09255首次发表:更新:

AI 中文总结

本研究对比RECS与ResStock数据集,发现树集成模型在住宅能源估算中,全特征性能优异,有限输入下精度收敛,同质队列建模可提升精度,需关注特征与数据集来源。

AI 中文摘要

在缺乏详细围护结构特性、设备效率、渗透数据、传感器或账单数据时,往往需要住宅能源估算。本研究利用美国两个全国代表性住宅能源数据集(基于调查的住宅能源消耗调查RECS和基于模拟的ResStock数据集)量化预测精度与输入可及性之间的权衡。首先使用全特征模型建立数据集特定的性能基准,针对总能源估算,随后将模型限制为10个低负担变量,这些变量可从居住者、行政记录或基于位置的天气数据中获取,无需现场能源审计。在CatBoost、XGBoost、LightGBM、随机森林和神经网络中,CatBoost在全特征分析中始终实现最高预测性能,ResStock的R²达0.90,RECS的R²达0.73。当特征集限制为10个房主可获取的输入以模拟现实部署条件时,模型性能收敛至RECS的R²=0.61、ResStock的R²=0.62,表明算法复杂性无法完全补偿缺失的物理和行为信息。但对于更同质的ResStock队列(包括2000-2010年建造、位于6A气候区、采用天然气供暖的独立单户住宅),减少输入的模型将精度提升至R²=0.85,证明针对同质人群进行建模的价值。结果表明,树集成模型可作为全国规模住宅能源数据集的高保真模拟器,但需仔细考虑特征可用性、数据集来源(经验vs合成)及适用用例。

英文摘要

Residential energy estimates are often needed before detailed envelope characteristics, equipment efficiencies, infiltration, sensor, or billing data are available. This study quantifies the trade-off between predictive accuracy and input accessibility using two nationally representative U.S. residential-energy datasets: the survey-based Residential Energy Consumption Survey (RECS) and the simulation-based ResStock dataset. Full-feature models were first used to establish dataset-specific performance benchmarks. For total-energy estimation, the models were subsequently restricted to ten low-burden variables obtainable from occupants, administrative records, or location-based weather data without an on-site energy audit. Among CatBoost, XGBoost, LightGBM, Random Forest, and Neural Networks, CatBoost consistently achieved the highest predictive performance for the full-feature analysis, reaching R2 = 0.90 for ResStock and R2 = 0.73 for RECS. When the feature set was restricted to ten homeowner-accessible inputs to simulate realistic deployment conditions, model performance converged to R2 = 0.61 for RECS and R2 = 0.62 for ResStock, showing that algorithmic complexity cannot fully compensate for missing physical and behavioral information. However, for a more homogeneous ResStock cohort consisting of single-family detached, natural-gas-heated homes in Climate Zone 6A constructed between 2000 and 2010, a reduced-input model improved accuracy to R2 = 0.85, demonstrating the value of targeted modeling for homogeneous populations. The results indicate that tree-based ensemble models can serve as high-fidelity emulators of national-scale residential energy datasets. However, careful consideration of feature availability, dataset origin (empirical vs. synthetic), and applicable use cases are also important.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑