Machine learning can support earlier estimates of hospital length of stay while giving greater attention to uncommon prolonged admissions. An analysis published in Healthcare Analytics tested a cost-sensitive prediction framework using administrative data from four clinical units in three hospitals in northern France. The models relied only on information available at admission and compared several ensemble methods with weighted versions designed to reduce the influence of highly imbalanced stay patterns. The strongest approach maintained high overall accuracy and improved performance for longer stays, although results varied between adult, paediatric and neonatal care.
Admission Data Support Early Prediction
The framework used 16,590 inpatient episodes from cardiology, general medicine, paediatrics and neonatology. The selected units differed in patient populations, organisation and typical stay patterns, creating a broad test of performance across hospital settings. General medicine had longer and more variable stays, while paediatric and neonatal stays were shorter and more concentrated.
Must Read: Human-AI Co-Design Refines Clinical Prediction Models
Only data available when the patient entered hospital were included. These covered age, diagnosis information, admission type, previous hospital use, clinical unit and selected contextual factors. Additional variables reflected previous stay duration and average stay patterns for diagnoses and units. The final feature set was chosen to balance predictive value, interpretability and overlap between variables. Records with inconsistent or implausible information were removed, while remaining missing values were handled using statistics derived from the training data to limit leakage.
The data were divided into development and held-out testing groups while preserving the distribution of units. Random Forest, Gradient Boosting and XGBoost were evaluated as continuous prediction models. The weighted versions of XGBoost gave greater importance to rarer stay durations during training. Several weighting approaches and intensities were tested to determine whether additional attention to prolonged admissions could improve prediction without weakening accuracy across the wider population. Performance was assessed globally, within each unit and among patients with longer stays.
Weighted Models Improve Longer-Stay Predictions
XGBoost achieved the strongest overall results. Its predictions differed from observed stay duration by an average of about 0.8 days and most were within one day. Random Forest and Gradient Boosting produced larger errors and explained less variation in stay length, indicating weaker performance on the heterogeneous administrative data.
Cost-sensitive weighting did not materially change overall average error. Across the different settings, global results remained close to those of standard XGBoost and confidence intervals overlapped. Some configurations produced statistically detectable changes in the distribution of individual errors, but these differences were small and did not represent meaningful gains in overall accuracy. The principal effect of weighting was therefore a redistribution of errors rather than a broad improvement across all admissions.
The advantage was clearer for stays of at least eight days. All weighted XGBoost configurations reduced error in this subgroup compared with the standard model and more than nine in ten predictions were within one day for most settings. A square-root weighting approach at moderate intensity produced the lowest error, while a logarithmic approach at mild intensity delivered a closely comparable result. Stronger weighting did not consistently perform better. The findings show that cost-sensitive training can direct model attention towards uncommon, operationally important stays while preserving stable performance across the full test population.
Clinical Units Show Different Levels of Predictability
Prediction quality differed substantially across the four units. Cardiology and general medicine were the most predictable, with high agreement between expected and observed stay duration. Their greater variation in length of stay and more standardised care pathways provided stronger signals from information available at admission.
Paediatric and neonatal care were more difficult to model. Paediatric trajectories can change rapidly and may depend on information that is not yet available when the patient arrives. Neonatal stays were tightly concentrated, so even relatively small errors produced weak values on measures that depend on outcome variation. Standard XGBoost still achieved an average error below one day in both settings, but its fit was less consistent than in the adult units.
Cost-sensitive weighting produced modest improvements in paediatrics and neonatology, including lower average errors and more stable performance measures. These gains were smaller than the difference between XGBoost and the other ensemble methods, but they indicate that weighting may be most useful where prediction is affected by greater uncertainty or limited variation in the available data.
The framework is positioned as a decision-support input rather than a standalone operational trigger. In adult units, early estimates may contribute to bed and staffing planning. In paediatric and neonatal care, they may be more appropriate as early indicators of risk than as precise forecasts. Weighting can also be adjusted to reflect local priorities, although mild or moderate settings may be preferable to more aggressive rebalancing.
Cost-sensitive machine learning can improve attention to prolonged hospital stays without reducing overall predictive performance. XGBoost delivered the strongest results across the full dataset and remained stable when weighting was introduced. The clearest gains appeared among longer admissions and in units where admission-level prediction was more difficult. The findings are limited to four units in three French hospitals and to records collected between 2017 and 2019. Broader validation and the addition of information gathered during the hospital stay are needed, particularly for paediatric and neonatal care.
Source: Healthcare Analytics
Image Credit: iStock