发表机构
Jatiya Kabi Kazi Nazrul Islam University(贾蒂亚·卡比·卡齐·纳兹鲁尔伊斯兰大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出一种机器学习框架,利用运营与情境特征预测餐厅每日食物浪费量,集成方法优于线性模型,随机森林表现最佳,并公开数据与代码以支持可持续管理。
AI 中文摘要
餐厅部门的食物浪费对环境可持续性和经济效率构成了重大挑战。本文提出了一种探索性的机器学习框架,用于根据运营和情境特征估算餐厅每日食物浪费量。通过整合餐厅需求记录、气象数据和时态事件指标,构建了一个结构化数据集,共包含27个特征下的77,980条记录。由于大规模的真实食物浪费测量数据无法公开获取,目标变量基于运营上合理的假设推导得出,并公开了完整的构建公式和受控的随机变异性,以确保完全可复现。在尊重餐厅运营时间顺序的按时间顺序70-30训练-测试划分下,评估了四种监督回归模型,即线性回归、决策树、随机森林和梯度提升,并辅以5折时间序列交叉验证。所有报告的指标明确限定为针对构建目标的性能,并不暗示对实测食物浪费的验证。集成方法始终优于线性基线。在系统排除代数上易泄漏的变量后,随机森林在现实特征子集上取得了MAE为6.19千克、RMSE为8.36千克和R²为0.817的成绩。特征重要性分析确定了菜单多样性、运营区域和时态活动模式为主要预测驱动因素。完整数据集、目标构建公式、代码库和实验配置已公开发布,以支持可复现性和未来扩展到经验测量的浪费数据。
英文摘要
Food waste in the restaurant sector poses a substantial challenge to environmental sustainability and economic efficiency. This paper presents an exploratory machine learning framework for estimating daily restaurant food waste quantities from operational and contextual features. A structured dataset was constructed by integrating restaurant demand records, meteorological data and temporal event indicators, yielding 77,980 records across 27 features. Because large-scale ground-truth food waste measurements are not publicly available, the target variable was derived from operationally justified assumptions, with the complete construction formula and controlled stochastic variability disclosed for full reproducibility. Four supervised regression models, namely Linear Regression, Decision Tree, Random Forest and Gradient Boosting, were evaluated under a chronological 70-30 train-test split that respects the temporal ordering of restaurant operations, augmented by 5-fold time-series cross-validation. All reported metrics are explicitly scoped to performance against the constructed target and do not imply validation against measured food waste. Ensemble methods consistently outperformed linear baselines. Random Forest attained an MAE of 6.19 kg, RMSE of 8.36 kg and $R^2$ of 0.817 on the realistic feature subset following systematic exclusion of algebraically leakage-prone variables. Feature importance analysis identified menu diversity, operational area and temporal activity patterns as the primary predictive drivers. The full dataset, target construction formula, codebase and experimental configurations are publicly released to support reproducibility and future extension to empirically measured waste data.
Comments6 pages, 6 figures. Accepted and presented at the 2026 IEEE International Conference on Signal Processing, Information, Communication and Systems (SPICSCON 2026). Accepted for publication in the conference proceedings