arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

决策模型能在何处诊断暖通空调故障?推理需求、物理表示与偏移下的鲁棒性

Where Can a Decision Model Diagnose HVAC Faults? Reasoning Demand, Physical Representation, and Robustness Under Shift

Wooyoung Jung

arXiv 2610.09937首次发表:更新:

发表机构

University of Arizona(亚利桑那大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究探究决策模型在暖通空调故障诊断中的能力边界,通过128个故障日测试,发现其在物理特征输入下能诊断单特征证据的故障,但需上下文,且在数据偏移下保持鲁棒性,支持代码计算物理、模型排序候选故障的分工。

AI 中文摘要

人工智能以多种形式支持建筑运维,每种形式都有其自身的障碍。专家规则必须针对每个系统进行调优,监督模型需要建筑很少记录的有标签数据,而语言模型返回的自由文本需要人工介入检查,因为它们所声明的置信度并不可靠。一种较新型的预训练模型,此处称为决策模型,为每个允许的答案返回一个概率,因此一个模型无需训练即可服务于多个决策。本研究回答了供暖、通风与空调系统故障诊断中的三个开放性问题:此类模型能做出哪些决策,需要什么输入,以及当条件变化时其概率是否成立。在来自四个真实设备公共数据集的128个故障日中,故障根据其诊断所需的推理进行分级,数据以原始形式、物理特征形式或带有Brick拓扑的形式给出。决策模型Jev、开放语言模型和一个监督模型面临九项测试,这些测试改变季节、控制配置或建筑。在给定物理特征的情况下,Jev和较大的开放模型诊断了其证据由单一特征承载的故障,但未能诊断需要运行上下文的故障。在偏移下,它们保持了准确性和校准,而监督模型损失了0.33的宏F1分数,但在建筑内领先或持平。它们的概率仍需要校正,且检测能力较弱。本研究绘制了决策模型能诊断哪些故障以及从什么输入进行诊断的图谱,并支持一种分工,即代码计算物理量,模型为操作员对候选故障进行排序。

英文摘要

Artificial intelligence supports building operations in several forms, each with its own barrier. Expert rules must be tuned for every system, supervised models need labeled data that buildings rarely record, and language models return free text that requires human-in-the-loop checking, since their stated confidence is unreliable. A newer kind of pretrained model, here called a decision model, returns a probability for every allowed answer, so one model could serve many decisions without training. This study answers three open questions for fault diagnosis in heating, ventilation, and air-conditioning systems: which decisions such a model can make, what input it needs, and whether its probabilities hold when conditions change. On 128 fault days from four public datasets of real equipment, faults are graded by the reasoning their diagnosis demands, with data given raw, as physical features, or with Brick topology. The decision model Jev, open language models, and a supervised model face nine tests that change season, control configuration, or building. Given physical features, Jev and the larger open model diagnosed faults whose evidence one feature carries, but not faults that need operating context. Under shift they kept their accuracy and calibration, while the supervised model lost 0.33 macro-F1 yet led or tied within a building. Their probabilities still needed correction, and detection was weak. The study maps which faults a decision model can diagnose and from what input, and supports a division of work in which code computes the physics and the model ranks candidate faults for an operator.

Comments58 pages, 4 figures, 12 tables. Submitted to Energy and Buildings

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑