发表机构
Spark Works Ltd.; Industrial Systems Institute, Athena Research Center; Sapienza University of Rome(Spark Works有限公司; 雅典研究中心工业系统研究所; 罗马第一大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文介绍Tethys数据集,包含24栋建筑91个月的每小时用水量数据,并公开了数据质量细节、修正聚合误差的方法及处理流程,以支持需求预测和泄漏检测研究。
AI 中文摘要
水资源需求预测和泄漏检测方法依赖于公共数据集,而对于水资源领域,这些数据集稀缺、时间跨度短,或仅在未经记录的清洗过程后发布,从而掩盖了其来源部署中的缺陷。我们提出了Tethys数据集:来自市政供水网络中24栋建筑的91个月每小时用水量数据,并随数据发布了对其质量的定量说明,而非以质量说明替代数据本身。原始可用率为59.4%,而数据缺失并非随机:4次全网络范围的故障共持续595天,导致整个设施同时中断。我们发现,生成发布文件时的聚合过程在累积指数中悄然引入了205,200次不可能出现的下降,而修正这一问题仅需一行代码的更改。由于水表是累积式的,短时间间隙两侧的读数可确定该间隙期间通过的水量,因此75.9%的小时数据基于实际测量,而24.1%的小时数据被报告为未知。我们发布了该数据集、其逐小时来源信息以及生成该数据的处理流程。
英文摘要
Methods for water demand forecasting and leak detection are based on public datasets, and for water those are scarce, short, or released only after an undocumented cleaning process, hiding defects of the deployment they came from. We present Tethys: 91 months of hourly water consumption data from 24 buildings of a municipal water network, published with a quantitative account of its quality, rather than in place of one. Raw availability is 59.4%, while loss is not random: 4 fleet-wide outages totalling 595 days interrupt the entire estate at once. We show that the aggregation producing the released files silently introduced 205,200 impossible decreases in a cumulative index, and that correcting it is a one-line change. Because the meters are cumulative, the readings bracketing a short gap fix the volume that passed through it, so 75.9% of hours rest on a measurement, while 24.1% are reported as unknown. We release the dataset, its per-hour provenance, and the pipeline that produces it.
CommentsPreprint submitted to and accepted for publication at the Workshop on Water Supply Systems of the Future, 12th IEEE International Smart Cities Conference 2026 (ISC2 2026)