Background: Postoperative pneumonia is a severe complication after hip fracture surgery and is associated with substantial morbidity and mortality. Conventional regression-based approaches have limitations in modeling complex interactions and perform ...
Background: Postoperative pneumonia is a severe complication after hip fracture surgery and is associated with substantial morbidity and mortality. Conventional regression-based approaches have limitations in modeling complex interactions and perform suboptimally in highly imbalanced datasets. Machine-learning techniques may offer improved predictive performance and facilitate early identification of high-risk patients. To develop and evaluate machine learning–based models for predicting postoperative pneumonia following hip fracture surgery and to identify key perioperative predictors using multiple interpretability methods.
Methods: This retrospective single-center study included 1,771 patients who underwent hip fracture surgery between 2005 and 2024. Postoperative pneumonia occurred in 66 patients (3.7%). Three machine learning algorithms-Extreme Gradient Boosting (XGBoost), Random Forest, and Support Vector Machine-along with a voting ensemble classifier, were trained using a 6:3:1 train-validation- test split. Four resampling strategies (SMOTE-NC, SMOTE-ENN, ADASYN, and Borderline-SMOTE) were applied to address class imbalance. Feature selection incorporated statistical screening, tree-based importance, permutation importance, SHAP, and an integrated importance score. Model performance was evaluated using area under the receiver operating characteristic curve (AUC), recall, precision, F1 score, and Matthews correlation coefficient.
Results: Integrated feature selection based on XGBoost feature importance, permutation importance, and SHAP values highlighted a set of perioperative predictors-age, body mass index, Charlson Comorbidity Index, operation time, intensive care unit admission, preoperative laboratory findings (hemoglobin, albumin, protein, and blood urea nitrogen), and postoperative hemoglobin. Using this reduced feature set, the voting ensemble achieved the highest overall discrimination across resampling strategies, with AUC values consistently ranging from 0.94 to 1.00. ADASYN produced the highest sensitivity, frequently achieving perfect recall by reducing false-negative predictions. Confusion matrix analysis showed that ADASYN and SMOTE-NC nearly eliminated false negatives, whereas SMOTE- ENN reduced false positives but showed decreased overall performance.
Conclusion: Machine learning models demonstrated strong potential for predicting postoperative pneumonia after hip fracture surgery, with a voting ensemble outperforming individual classifiers across evaluation metrics. ADASYN substantially improved sensitivity by reducing false-negative predictions in the setting of extreme class imbalance. These findings suggest that machine learning-based risk stratification using routinely available perioperative clinical and laboratory data may support earlier identification of high-risk patients for postoperative pneumonia and enhance perioperative management in geriatric hip fracture care.