Where Can a Decision Model Diagnose HVAC Faults? Reasoning Demand, Physical Representation, and Robustness Under Shift
Decision model Jev diagnoses simple HVAC faults from physical features but not context-heavy faults, and stays calibrated under shift.
The study tests whether a pretrained decision model, Jev, can diagnose HVAC faults without task-specific training, returning a probability for each allowed answer. On 128 fault days from four public datasets, inputs were raw data, physical features, or Brick topology, and faults were graded by reasoning demand. Given physical features, Jev and a larger open language model diagnosed faults carried by one feature but not faults needing operating context. Under shifts of season, control configuration, or building, they kept accuracy and calibration, while a supervised model lost 0.33 macro-F1 yet led or tied within a building; probabilities still needed correction and detection was weak.
- Evaluated on 128 fault days from four public equipment datasets.
- With physical features, Jev diagnosed single-feature faults but not context-heavy faults.
- Under season, control, and building shift, Jev kept accuracy and calibration.
- A supervised model lost 0.33 macro-F1 under shift but led within a building.
- Probabilities still needed correction, and fault detection remained weak.
Full article248 words · extracted from arxiv.org · click to collapse
Artificial intelligence supports building operations in several forms, each with its own barrier. Expert rules must be tuned for every system, supervised models need labeled data that buildings rarely record, and language models return free text that requires human-in-the-loop checking, since their stated confidence is unreliable. A newer kind of pretrained model, here called a decision model, returns a probability for every allowed answer, so one model could serve many decisions without training. This study answers three open questions for fault diagnosis in heating, ventilation, and air-conditioning systems: which decisions such a model can make, what input it needs, and whether its probabilities hold when conditions change. On 128 fault days from four public datasets of real equipment, faults are graded by the reasoning their diagnosis demands, with data given raw, as physical features, or with Brick topology. The decision model Jev, open language models, and a supervised model face nine tests that change season, control configuration, or building. Given physical features, Jev and the larger open model diagnosed faults whose evidence one feature carries, but not faults that need operating context. Under shift they kept their accuracy and calibration, while the supervised model lost 0.33 macro-F1 yet led or tied within a building. Their probabilities still needed correction, and detection was weak. The study maps which faults a decision model can diagnose and from what input, and supports a division of work in which code computes the physics and the model ranks candidate faults for an operator.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2610.09937