A calibrated machine-learning approach may strengthen early warning for in-hospital cardiac arrest by combining underlying disease, laboratory results and changing physiology. The retrospective analysis, published in BMJ Health & Care Informatics, used records from adult general wards at a tertiary hospital in Taiwan between 2019 and 2021. It compared 115 patients who experienced cardiac arrest with 115 age-matched and ward-matched controls. Vital signs and consciousness were assessed in three time windows before the event, while four prediction models were trained, calibrated and tested on a temporally separate cohort. The strongest model supported high-sensitivity detection but still requires broader validation before clinical use.
Combining Baseline Risk with Evolving Physiology
The approach expanded conventional early warning beyond short-term vital signs. Demographic information, comorbidities and the most recent laboratory results were drawn from electronic health records. Physiological measurements were organised into periods covering 16–24 hours, 8–16 hours and 1–8 hours before cardiac arrest or the equivalent index time for controls. When several readings were available, the most abnormal value within each period was used.
The Modified Early Warning Score combined temperature, heart rate, respiratory rate, systolic blood pressure, oxygen saturation, supplemental oxygen use and level of consciousness. Its association with cardiac arrest was strongest in the earlier windows. Each one-point increase was linked to higher odds of arrest during both the 16–24-hour and 8–16-hour periods. In the final eight hours, lower oxygen saturation was the only physiological variable that remained associated with the event after adjustment. Each one percentage-point fall in oxygen saturation corresponded to higher odds of arrest.
This pattern does not establish that oxygen saturation should replace the broader score. Oxygenation is already included within the score and overlapping variables may affect the findings. Instead, the results indicate that the combined score may capture broader deterioration earlier, while oxygen desaturation becomes a stronger marker as arrest approaches. The time-based pattern remains exploratory and requires confirmation in larger prospective patient groups.
Comorbidities and Laboratory Results Add Predictive Signal
Several baseline characteristics were independently associated with cardiac arrest. Atrial fibrillation showed the strongest association, followed by heart failure and end-stage renal disease. Male sex was also linked to higher odds of an event. Age and length of stay did not differ significantly between the matched groups, although age matching limited assessment of age-related differences.
Must Read: Explainable AI Visualisations for ICU Intubation Risk
Routine laboratory measures added further information. Higher potassium concentration and white blood cell count were independently associated with cardiac arrest. The mean potassium level was only modestly higher among patients who experienced arrest and values in both groups remained within the normal reference range. White blood cell counts were also higher in the arrest group. Sodium, calcium and glucose did not differ significantly.
The findings support the use of comorbidities and laboratory indices alongside physiological observations rather than relying exclusively on threshold-based early warning scores. However, the associations should be treated as risk markers rather than causal factors. The retrospective design cannot exclude unmeasured differences in illness severity, treatment limits or monitoring intensity.
Patients with do-not-resuscitate orders were excluded, and the population was limited to adult medical and surgical general wards. Emergency departments, intensive care units, obstetric and paediatric services, psychiatry and long-term care settings were not included. These selection criteria define the population represented by the model and limit transfer of the results to other clinical environments.
Calibration Strengthens Probabilities but Deployment Needs Testing
Logistic regression, random forest, histogram-based gradient boosting and XGBoost models were trained using fivefold cross-validation. Each matched pair remained within the same training fold to reduce information leakage. Records from 2019 and 2020 supported model development and calibration, while matched pairs from 2021 formed a separate temporal test group. Missing values were imputed using models fitted only on training data.
All four models underwent isotonic probability calibration. This step aimed to align predicted probabilities more closely with observed outcomes within the matched dataset. Calibrated XGBoost produced the best overall performance, with strong discrimination and average precision in temporal testing. At a sensitivity close to 95%, specificity was approximately 45% to 50%, reflecting an operating point designed to minimise missed cardiac arrests.
Model comparisons remain exploratory because the dataset was modest and came from one hospital. The matched case–control design also created an event prevalence far higher than would occur across a full ward population. As a result, positive predictive value, precision and absolute probabilities cannot be transferred directly into routine practice. Recalibration to local event prevalence would be necessary before deployment.
Threshold selection would also need to account for local resources, acceptable alert burden and the response pathway attached to each alert. Prospective evaluation should determine whether calibrated warnings improve care when integrated into standardised ward surveillance and rapid-response processes.
Combining comorbidities, routine laboratory measures and time-windowed physiological data produced a high-sensitivity warning approach for cardiac arrest on general wards. Earlier deterioration was reflected by the Modified Early Warning Score, while oxygen desaturation became more prominent during the final eight hours. Calibrated XGBoost performed best in temporal testing, but its probabilities reflect the matched dataset rather than routine ward prevalence. The single-centre setting, retrospective design and modest sample require cautious interpretation. Recalibration and prospective multicentre validation remain necessary before the model can support clinical implementation.
Source: BMJ Health & Care Informatics
Image Credit: iStock
References:
Yu W, Pan M & Chen C (2026) Prediction of in-hospital cardiac arrest on general wards using calibrated machine learning. BMJ Health & Care Informatics;33:e101893.