Open accessFrontiers in Medicine · 1 June 2026

Comparing Manual vs. Automated Machine Learning and Deep Learning Models for Predicting One-Year Mortality in Elderly Hip Fracture Patients

Adi Shuchami, Maxim Glebov, Maksim Katsin, Yotam Portnoy, Haim Berkenstadt, Dina Orkin, Teddy Lazebnik

Affiliations
  1. Department of Mathematics, Ariel University, Ariel, Israel
  2. Department of Anesthesiology, Sheba Medical Center, Ramat Gan, Israel
  3. Department of Information Systems, University of Haifa, Haifa, Israel
  4. Faculty of Medicine, Tel-Aviv University, Tel Aviv, Israel
  5. Department of Computing, Jönköping University, Jönköping, Sweden

The paper at a glance

Hip fractures carry a high risk of death in older patients, and accurate risk prediction can help plan their care. Using data from 2,604 patients aged 65 or older who had urgent hip fracture surgery at Sheba Medical Center, we compared machine learning and deep learning models for predicting death within one year. A manually tuned gradient boosting model (XGB) performed best, with an AUC of 0.846, and an automated pipeline built with a large language model performed comparably, with an AUC of 0.844.

2,604elderly hip fracture surgery patients
0.846AUC of the best, manually optimized XGB model
0.844AUC of the LLM-generated automated model

Key findings

  • Manually optimized Extreme Gradient Boosting performed best (AUC=0.846, accuracy=0.791, F1-score=0.667).
  • Important predictors included baseline serum albumin and urea levels, patient age, hypothermia during surgery and the number of chronic diseases.
  • An automated pipeline generated with a large language model and TPOT matched XGB closely (AUC=0.844), with higher recall but slightly lower precision.
FIGURE 3 The ROC curves illustrate the performance of all ten models in the training cohort. Shaded areas represent the 95% confidence intervals derived from the k-fold (k = 5) cross-validation splits. (a) Training set; (b) Test set.
FIGURE 3 The ROC curves illustrate the performance of all ten models in the training cohort. Shaded areas represent the 95% confidence intervals derived from the k-fold (k = 5) cross-validation splits. (a) Training set; (b) Test set. See it in the paper
On this page
  1. Abstract
  2. Introduction
  3. Materials and methods
  4. Ethics approval and reporting guidelines
  5. Study design and population
  6. Data collection
  7. Sample size calculation
  8. Outcome
  9. Results
  10. Descriptive statistics
  11. Model performance
  12. Feature importance
  13. Automatic ML model
  14. Discussion
  15. Data availability statement
  16. Ethics statement
  17. Author contributions
  18. Funding
  19. Acknowledgments
  20. Conflict of interest
  21. Supplementary material
  22. Article notes
  23. References

Abstract

Background: Hip fractures are associated with significant mortality, especially among elderly patients. Accurate prediction of mortality risk is crucial for optimising perioperative care and resource allocation. Recent advances in machine learning (ML) and deep learning (DL) offer promising methods to enhance clinical risk prediction models; however, their clinical implementation often remains limited due to the complexity of these techniques. Methods: This retrospective cohort study included 2,604 elderly patients (≥65 years) undergoing urgent hip fracture surgery at Sheba Medical Center, Israel, between January 2017 and November 2023. Multiple ML and DL algorithms were evaluated for predicting one-year all-cause mortality using a comprehensive set of clinical, demographic, perioperative, and laboratory variables. Models were rigorously developed and validated through stratified 5-fold cross-validation, addressing class imbalance with the Synthetic Minority Oversampling Technique (SMOTE). Additionally, an automated ML pipeline, generated using a large language model (LLM) coupled with the Tree-based Pipeline Optimisation Tool (TPOT), was benchmarked against manually optimised models. Model performances were assessed using area under the receiver operating characteristic curve (AUC), accuracy, precision, recall, F1-score, false-positive rate, and true-negative rate, supplemented by permutation importance and SHapley Additive exPlanations (SHAP) for interpretability. Results: Among all models evaluated, the manually optimised Extreme Gradient Boosting (XGB) algorithm demonstrated superior predictive performance (AUC=0.846, accuracy=0.791, F1-score=0.667, precision=0.773, NPV=0.798). Important predictors identified included baseline serum albumin and urea levels, patient age, intraoperative hypothermia, and the number of chronic diseases. The automated ML model, generated via LLM and TPOT frameworks, showed comparable performance to the XGB model (AUC=0.844), with a higher recall but slightly lower precision. Discussion: ML-based models, particularly the XGB algorithm, significantly enhance predictive accuracy for one-year mortality among elderly hip fracture patients. Crucially, an automated ML framework leveraging large language models provides a practical, clinically accessible alternative, effectively democratising advanced predictive analytics in healthcare settings.

Introduction

Hip fracture represents a significant global health concern, with projections estimating approximately 4.5 million annual cases worldwide by 2050 (1). This injury predominantly affects older adults, leading to substantial clinical and economic burdens on healthcare systems globally (2). Postoperative mortality rates remain alarmingly high, with studies reporting up to 36% mortality within the first year following surgical intervention (3). These concerning figures highlight an urgent need for accurate risk stratification tools capable of identifying high-risk patients and guiding perioperative management, ultimately aiming to reduce preventable mortality.

Advancements in data-driven approaches, particularly machine learning (ML), have enabled the development of sophisticated clinical prediction models integrating perioperative and broader clinical variables (4). Unlike traditional regression techniques, ML algorithms effectively manage real-world data characterized by complex, nonlinear predictor interactions, often resulting in superior predictive accuracy (5). Indeed, comparative studies assessing ML-based models against conventional statistical approaches in predicting hip fracture mortality have consistently demonstrated enhanced performance, reinforcing the rationale for further exploration and adoption of ML methodologies (6–8).

This study aims to evaluate various ML and deep learning (DL) algorithms for predicting one-year mortality among elderly hip fracture patients. Crucially, we investigate the effectiveness of an automated ML platform facilitated by a large language model (LLM) in empowering clinicians with limited technical expertise to independently develop robust and clinically meaningful prediction models. By specifically addressing this practical aspect, our study underscores the potential for automated ML platforms to significantly democratize access to advanced predictive analytics in clinical settings. We aim to benchmark the performance of these automated ML and DL approaches against established predictive models, thereby demonstrating the added value and practicality of such frameworks for enhancing clinical risk prediction and informed decision-making.

Materials and methods

Ethics approval and reporting guidelines

This study received ethical approval from the Institutional Ethics Committee at Sheba Medical Center, Israel (Approval No. SMC-D 0976-24), and was performed in accordance with the relevant guidelines and regulations in accordance with the Declaration of Helsinki. Due to the retrospective observational design of the study, the requirement for informed consent was waived by the committee. The study protocol and reporting follow the Transparent Reporting of a Multivariable Prediction Model for Individual Prognosis or Diagnosis with Artificial Intelligence (TRIPOD-AI) guidelines.

Study design and population

This retrospective cohort study was conducted at a single tertiary care center (Sheba Medical Center, Israel). The study population included elderly patients, aged 65 years and older, undergoing primary emergency surgery for hip fracture between 1 January 2017 and 1 November 2023. Data extraction was performed using electronic health records from the institution’s database system (Chameleon, Elad Software Ltd., Tel Aviv, Israel). Eligible cases were identified through International Classification of Diseases, Tenth Revision (ICD-10) diagnosis codes: S72.0 (Fracture of neck of femur), S72.1 (Pertrochanteric fracture), and S72.2 (Subtrochanteric fracture). Patients were included if they had complete clinical data and a minimum follow-up of 1 year or until death. Exclusion criteria encompassed high-energy trauma mechanisms, multiple fractures, open fractures, and pathological fractures secondary to malignancy or metabolic bone disease.

Data collection

Data collected encompassed a comprehensive set of variables, including demographic characteristics, functional and social assessments, comorbidities, laboratory parameters, physiological measurements, intraoperative details, and postoperative outcomes. A detailed description of variables is provided in Supplementary Table 1.

The dataset was extracted from the hospital electronic medical record (EMR) environment, where data were already organized within a structured clinical data framework designed to support harmonization, formatting consistency, and routine data quality control. In accordance with best practices for clinical ML preprocessing, the dataset was reviewed to confirm variable definitions, preserve clinically appropriate coding, and ensure that the exported table was analysis-ready for supervised learning. Because the data originated from an EMR system with standardized organization, major preprocessing focused primarily on verification of completeness, consistency, and correct variable representation rather than large-scale manual restructuring. Subject-matter expertise remained essential at this stage to confirm the clinical meaning of variables, ensure that preprocessing decisions were medically appropriate, and verify that the final dataset accurately reflected real-world perioperative and baseline patient information. These preprocessing steps are important for both modeling strategies presented in this study to ensure clinical relevance and usability.

Sample size calculation

Given the relatively balanced nature of our dataset, we included a 20% contingency in our initial sample size estimation to account for potential undersampling and to maintain data balance. Considering the complexity of the predictive model and the total number of features, we adhered to the established methodological recommendation of at least 10 events per variable (EPV) (9). Since our dataset includes 98 variables, the minimum sample size required is 1,176 patients.

Outcome

The primary outcome was all-cause one-year mortality, defined as death occurring within 1 year from the date of fracture. Mortality data were obtained from the Ministry of Health’s vital status registry, and accuracy was subsequently validated by manual review of patient medical records.

Model development: manual model development and evaluation

Figure 1 illustrates a schematic representation of the methodological framework employed in this study.

Initially, descriptive statistical analyses were conducted to characterize the dataset. Subsequently, a Pearson correlation matrix was generated to investigate linear associations among features and their correlation with one-year mortality.

To develop machine learning (ML) and deep learning (DL) predictive models, the dataset was partitioned into training (80%) and validation cohorts (20%). The training cohort was further subjected to k-fold cross-validation (10) to ensure robust model evaluation. In this process, the data were divided into five equally sized, mutually exclusive subsets. Each subset served as the validation cohort exactly once, while the remaining subsets formed the training set, resulting in five validation iterations. Model performance was averaged across iterations to estimate predictive capability reliably.

Both training and validation cohorts were stratified by age and sex to ensure representative distributions. Within each cross-validation fold, an optimisation process was employed to minimise demographic differences in age and sex distributions across subsets. This optimisation, analogous to the nurse scheduling problem (11)—an NP-hard computational problem (12)—was addressed using Directed Bee Colony Optimisation to achieve a near-optimal solution (13).

Following data partitioning, seven ML models and three DL models were developed and comparatively evaluated. ML algorithms included logistic regression (LR) (14), naïve Bayes (NB) (15), k-nearest neighbours (kNN) (16), support vector machine (SVM) (17), decision tree (DT) (18), random forest (RF) (19), and XGBoost (XGB) with 2 L regularization (20). DL algorithms comprised multilayer perceptron (MLP) (21), TabNet (22), and a transformer-based tabular model (TabTransformer, TT) (23). Hyperparameter tuning was performed via grid search (24), systematically testing parameter combinations to optimise model accuracy during cross-validation.

Because the primary outcome was moderately imbalanced (one-year mortality rate: 36.9%), we evaluated each model under two settings: non-balanced, in which models were trained using the original class distribution, and balanced, in which Synthetic Minority Oversampling Technique (SMOTE) (25) was applied to the training folds to generate a more even class distribution. The validation/test data were kept unchanged to preserve a realistic clinical evaluation setting. This comparison allowed us to assess how class balancing influences predictive performance, particularly the trade-off between sensitivity to mortality events and false-positive predictions. This comparison provided insights into the impact of data balancing on model accuracy and feature importance.

Model performance in the validation cohort was evaluated comprehensively using multiple metrics: accuracy, F1-score, recall (sensitivity), precision, false positive rate (FPR), true negative rate (TNR), negative predictive values (NPV), and the area under the receiver operating characteristic (ROC) curve (AUC). Feature importance for each model was assessed using permutation importance (26), quantifying how random permutation of each feature influenced model performance. Additionally, SHapley Additive exPlanations (SHAP) (27) were utilised to offer detailed, instance-level insights into individual feature contributions to model predictions. All analyses are performed using the Python programming language.

Model development: automatic model development and evaluation

As the analytical procedures outlined above require advanced expertise in ML and DL, or at least significant experience in data-driven model development, they remain challenging and largely inaccessible for most clinicians. However, recent advancements in automated ML frameworks, coupled with the wider availability of LLMs such as ChatGPT (28), have prompted exploration into whether these tools could facilitate the creation of “out-of-the-box” ML models. Specifically, we aimed to determine whether such automated approaches could generate a viable predictive model, evaluate its performance rigorously, and present results clearly and understandably to clinical professionals without specialized ML knowledge.

To investigate this, we provided the ChatGPT-o3 model (29) with a structured prompt (Supplemental material A). The model-generated Python code was executed directly, without modification (“as-is”), in a Google Colab environment, and the results obtained from this automated approach are presented below.

The structured prompt was carefully designed to replicate the analytical methodology applied previously in our manually developed models. For the automated ML component, we chose the Tree-based Pipeline Optimization Tool (TPOT) (30) due to its extensive prior application (31) and proven reliability in clinical prediction tasks (32). The prompt included a definition of the LLM’s “persona”—a method previously demonstrated to enhance LLM performance (33)—and concluded with explicit, directive instructions detailing the desired analytical outcomes.

A schematic view of the methodological framework of this study
FIGURE 1 A schematic view of the methodological framework of this study.

Predictive performance of the automated ML-generated model was rigorously assessed using a comprehensive set of metrics identical to those employed for manual model evaluations. Additionally, we analysed feature importance and applied SHAP to achieve an in-depth, instance-level interpretation of each feature’s contribution to the model’s predictions.

Results

Descriptive statistics

The final cohort included 2,604 patients, of whom 1719 (66.0%) were female. The mean patient age was 82.3 years (SD 8.3). Detailed demographic and baseline characteristics of the cohort are summarised in Table 1.

Figure 2 displays the Pearson correlation matrix for the variables incorporated into the predictive model and the target outcomes. Most predictors showed minimal intercorrelation, with correlation coefficients approaching zero. More importantly, as indicated by the last row, the source features are poorly linearly correlated to the target variable. Therefore, a nonlinear machine learning method is anticipated to demonstrate superior predictive performance compared with traditional linear approaches such as logistic regression. Notably, due to the symmetric nature of correlations, the values above and below the main diagonal line are identical.

Model performance

Figure 3 demonstrates the receiver operating characteristic (ROC) curves for all ten evaluated models. The results from the training cohort are presented with corresponding 95% confidence intervals (CIs), calculated using five-fold cross-validation. Overall, the models exhibited comparable discriminatory abilities, although the K-Nearest Neighbour (KNN), TabNet, and Naïve Bayes (NB) algorithms consistently demonstrated lower area under the curve (AUC) values compared to the remaining models across both training and testing datasets. Of particular note, the TT model exhibited the lowest performance, with an AUC of 0.580, suggesting possible underfitting. To provide a more comprehensive evaluation, we further analysed additional relevant performance metrics. Importantly, the figure presents the ROC curves for models trained on the original, non-resampled dataset, in order to compare discrimination under the observed class distribution.

TABLE 1 Demographic and baseline characteristics of the cohort.
CharacteristicAll cohort
Age (years)82.3 + −8.26
Female66.0%
BMI (kg/m2)25.08 + −4.29
Surgery duration (min)55 + −31
PACU time (min)108 + −52
Previous hospitalization in 6 months25.6%
Fracture to surgery time less than 48 h87.1%

All data are presented as the mean [SD] or n (%). BMI, Body mass index; PACU, Post-anesthesia care unit.

Table 2 summarises the performance metrics of ML and DL models on the test dataset, including area under the receiver operating characteristic curve (AUC), accuracy, F1 score, recall, precision, false-positive rate (FPR), true-negative rate (TNR), and negative predictive value (NPV). In the slightly imbalanced scenario, with a one-year mortality rate of 36.9%, the extreme gradient boosting (XGB) model demonstrated superior performance, achieving the highest AUC (0.846), accuracy (0.791), and F1 score (0.667). The Naive Bayes (NB) model exhibited the highest recall (0.968) but had the lowest accuracy (0.380). Although the DL-based TabTransformer (TT) model showed excellent precision (0.857), FPR (0.003), and TNR (0.997), its low recall (0.032) and modest AUC (0.580) suggest a bias towards predicting a single class, thus significantly restricting its clinical applicability.

In the balanced scenario, a similar performance pattern emerged, with XGB consistently outperforming other models across most metrics. However, the random forest (RF) model achieved a slightly higher AUC (0.849), and the NB model exhibited superior recall (0.962), surpassing XGB by absolute margins of 0.016 and 0.311, respectively. Nonetheless, NB continued to demonstrate substantially lower accuracy (0.376) compared to XGB (0.775). Overall, when considering all performance metrics comprehensively, the XGB model emerged as clearly superior and was consequently selected for subsequent analyses. Specifically, XGB is well-suited to the structured clinical dataset and also achieved the best F1 score, which considered a key metric because it captures the balance between recall and precision in an imbalanced mortality prediction task. From a clinical perspective, recall is especially important, since failing to identify a patient at high risk of one-year mortality may result in missed opportunities for closer follow-up, earlier intervention, or more intensive care planning. However, focusing on recall alone may favor models that label too many patients as high risk, thereby reducing precision. In practice, this can create a biased prediction pattern that overestimates risk and generates unnecessary clinical workload, including excess alerts, avoidable evaluations, and inefficient use of healthcare resources.

Feature importance

For interpretability analysis, we focused on the XGB model trained on the balanced dataset, as this version was selected for subsequent detailed analysis after comparison of balanced and non-balanced training. To this end, Figure 4 illustrates the distribution of feature importance derived from the XGB model trained on the balanced dataset. Due to the extensive number of variables, only the six most influential features are presented. The most significant predictors identified were baseline albumin, baseline urea, and patient age.

Figure 5 illustrates the SHAP analysis for the XGB model where each dot represents one individual patient/sample. The x-axis shows the SHAP value, indicating the direction and magnitude of that feature’s contribution to the predicted one-year mortality risk. Dot color reflects the original feature value (low to high). Features are ordered according to their overall importance in the model. This analysis identifies patient age as the most influential predictor, with older age strongly associated with an increased likelihood of one-year mortality. Baseline serum albumin, the second most significant feature, indicates that lower albumin levels correlate with elevated mortality risk. A comparable association is observed for intraoperative hypothermia, ranked as the third most impactful feature. Additionally, male sex is identified as a predictor associated with higher mortality rates. Finally, the American Society of Anesthesiologists (ASA) score demonstrates a monotonic relationship with the predicted probability of mortality within one year.

Pearson correlations between the dataset’s features
FIGURE 2 Pearson correlations between the dataset’s features. BMI, body mass index; ASA, American Society of Anesthesiologists Physical Status classification; num_dis, number of diseases; funcstat, functional status; cohabit, cohabitation status; mobility, mobility status; prevhosp., previous hospitalisation within 6 months; creat_base, baseline creatinine; urea_base, baseline urea; lact_base, baseline lactate; be_base, baseline base excess; hb_base, baseline haemoglobin; alb_base, baseline albumin; EF, ejection fraction; surg_min, surgery duration in minutes; hypoterm, intraoperative hypothermia; ioh55, intraoperative hypotension (MAP<55 mmHg); intrabld, intraoperative blood transfusion; preopbld, preoperative blood transfusion; pacu_min, duration in post-anaesthesia care unit in minutes.

Automatic ML model

In comparison, the automated ML framework, generated through code written by an LLM, identified an ensemble majority-vote model comprising CatBoost, kNN, and RF as the optimal solution for the given dataset. Specifically, the final prediction of this automated model reflects the consensus reached by at least two of the three constituent algorithms. The code for this automated framework, generated by the LLM, as well as the code for the model derived using the TPOT framework, is provided in Supplementary material B.

Table 3 presents a comparative analysis of the predictive performance between the automatically generated models (LLM and TPOT) and the manually optimized best-performing model (XGB). Overall, predictive performance metrics between the automated and manually developed models were comparable, with accuracy showing the largest discrepancy. Differences in area under the receiver operating characteristic curve (AUC) between both models were negligible in both balanced and unbalanced data scenarios (0.002 and 0.012, respectively). The manually developed XGB model exhibited superior precision, false positive rate (FPR), and true negative rate (TNR). Conversely, the automated ML model demonstrated superior performance in terms of the F1 score and recall metrics.

Figures 6a–c presents the ROC curve, feature importance, and SHAP analysis for the model automatically derived using the LLM and the automated machine learning framework (TPOT). The ROC curve demonstrates performance comparable to the XGB model. The feature importance analysis reveals substantial overlap with the top six features identified by the XGB model, albeit with minor variations in ranking. Similarly, SHAP analysis confirms this concordance, highlighting consistent clinical interpretability across both models.

Discussion

In this study, we systematically evaluated the potential of ML and DL models to support clinical decision-making by predicting one-year mortality in elderly patients with hip fractures. Our findings indicate that traditional ML methods generally outperform DL models when applied to structured tabular clinical data, a result consistent with recent comprehensive evaluations conducted across diverse clinical datasets (31, 34).

The ROC curves illustrate the performance of all ten models in the training cohort
FIGURE 3 The ROC curves illustrate the performance of all ten models in the training cohort. Shaded areas represent the 95% confidence intervals derived from the k-fold (k = 5) cross-validation splits. (a) Training set; (b) Test set.
TABLE 2 Performance of ML and DL models on the test dataset.
BalancedModelAUCAccuracyF1 scoreRecallPrecision
(PPV)
FPRTNRNPV
NoLR0.8380.7660.6280.5540.7250.1160.8840.781
NB0.7530.3800.5270.9680.3620.9460.0540.753
KNN0.7520.6930.4160.3060.6480.0930.9070.703
SVM0.8200.7500.6220.5750.6770.1520.8480.781
DT0.7500.6970.5030.4300.6060.1550.8450.728
RF0.8450.7540.5520.4250.7900.0630.9370.746
XGB0.8460.7910.6670.5860.7730.0960.9040.798
MLP0.8230.7580.6270.570.6970.1370.8630.782
Tabnet0.7070.6550.1180.0650.6670.0180.9820.655
TT0.5800.6530.0620.0320.8570.0030.9970.650
YesLR0.8270.7600.6940.7630.6370.2420.7580.828
NB0.7480.3760.5240.9620.3600.9490.0510.708
KNN0.6460.5970.5510.6940.4570.4570.5430.761
SVM0.7430.6780.5380.5270.5510.2390.7610.745
DT0.6010.6370.5380.5910.4930.3370.6630.741
RF0.8490.7720.6740.7200.6670.2000.8000.841
XGB0.8330.7750.6930.6510.6990.1550.8450.811
MLP0.7650.7180.6080.6130.6030.2240.7760.784
Tabnet0.7530.6580.6150.7630.5140.4000.6000.821
TT0.5030.4880.5570.9030.4030.7430.2570.826

ML, machine learning; DL, deep learning; LR, logistic regression; NB, naïve Bayes; KNN, k-nearest neighbours; SVM, support vector machine; DT, decision tree; RF, random forest; XGB, XGBoost; MLP, multi-layer perceptron; TT, TabTransformer; AUC, area under the receiver operating characteristic curve; FPR, false positive rate; TNR, true negative rate. Best-performing values for each metric are highlighted in bold. The balanced group indicates models trained using SMOTE-balanced data, while the other group indicates models trained on the original, non-resampled data.

Feature importance distribution for the XGBoost prediction model
FIGURE 4 Feature importance distribution for the XGBoost prediction model. alb_base, baseline albumin; hypoterm, intraoperative hypothermia; ASA, American Society of Anesthesiologists Physical Status classification; intraop phenylephrine, intraoperative use of phenylephrine.
SHAP analysis of the XGB model
FIGURE 5 SHAP analysis of the XGB model. alb_base, baseline albumin; hypoterm, intraoperative hypothermia; ASA, American Society of Anesthesiologists Physical Status classification; intraop phenylephrine, intraoperative use of phenylephrine.
TABLE 3 Performance comparison between the automatically derived model (LLM and TPOT) and the manually developed best-performing model (XGB).
BalancedModelAUCAccuracyF1 scoreRecallPrecisionFPRTNR
NoXGB0.8460.7910.6670.5860.7730.0960.904
LLM and TPOT0.8440.7260.6980.8870.5750.3640.636
YesXGB0.8330.7750.6930.6510.6990.1550.845
LLM and TPOT0.8250.7180.7200.8910.6050.1290.833

LLM, large language model; TPOT, Tree-based Pipeline Optimization Tool; XGB, XGBoost; AUC, area under the receiver operating characteristic curve; FPR, false positive rate; TNR, true negative rate. The best-performing values for each metric are highlighted in bold.

Among the various predictive models assessed, the manually developed XGB model demonstrated the highest predictive accuracy, thereby offering clinicians a robust and reliable tool to effectively identify high-risk patients. This finding aligns closely with previous clinical investigations utilising machine learning approaches, which consistently report superior performance of XGB in a variety of clinical prediction tasks due to its capacity to handle complex, high-dimensional data efficiently (35). Specifically, XGB was favored not only because of its overall predictive performance, but also because it achieved the best F1 score as indicated in Table 2. Such predictive accuracy has profound clinical implications, as early and precise identification of high-risk patients allows clinicians to implement timely interventions, potentially improving survival rates, reducing morbidity, and enhancing overall patient outcomes.

The strength of the XGB model primarily resides in its capacity to integrate multiple, heterogeneous clinical predictors, effectively mirroring the intricacies of real-world clinical practice. Key predictors identified by our analysis; such as patient age (36), baseline serum albumin (37), baseline serum urea (38), and intraoperative hypothermia (39), represent well-established clinical risk factors previously corroborated in the literature. This consistency significantly reinforces the clinical validity and interpretability of our findings (6). Furthermore, by quantifying the relative contributions of these predictors through advanced feature importance metrics and SHAP analyses, our study not only reinforces existing clinical knowledge but also provides clinicians with actionable insights into patient-specific risk profiles. Consequently, this facilitates targeted resource allocation, optimised care delivery, and prioritisation of interventions tailored to those patients identified as being at the highest risk.

Remarkably, we also explored the innovative approach of leveraging a large language model (ChatGPT-o3) to automatically generate ML code within the TPOT automated machine learning framework. The predictive performance of this automatically generated model was comparable, albeit slightly inferior, to that of the manually optimised XGB model. This similarity in both predictive performance and clinical interpretability underscores the promising utility of automated ML workflows, particularly for clinicians and institutions lacking extensive technical expertise in data science. Such an automated approach offers the distinct advantage of rapidly developing predictive analytics, facilitating continual model updating, refinement, and retraining as new clinical data becomes available, thus maintaining clinical applicability over time. Accordingly, the key difference between the two approaches is not the specific presence or absence of individual algorithms, but whether pipeline construction, model combination, and selection were performed manually by the research team or automatically by the LLM/TPOT workflow.

Importantly, in real clinical practice, both the manual ML workflow and the automated LAG/LLM-based workflow would require a human-in-the-loop. In the manual workflow, this would likely involve data scientists or clinicians with ML expertise performing dataset preparation, feature handling, model selection, tuning, and validation before the model could be implemented in a clinical decision-support setting. In the automated LLM-based workflow, expert oversight would still be necessary to formulate the clinical question, verify the suitability of the dataset, review the generated pipeline, and assess the clinical plausibility and safety of the resulting model (40). Thus, the difference between the two approaches is not the elimination of human involvement, but rather the degree of technical burden and specialization required. The automated approach may reduce the need for advanced hands-on ML expertise and shorten development time, whereas the manual workflow is generally more resource-intensive and dependent on dedicated methodological expertise. Accordingly, the automated workflow may be better viewed as a tool for reducing development barriers and resource demands, rather than as a fully autonomous replacement for expert-driven model development (41, 42).

(Continued)
FIGURE 6 (Continued)
Performance evaluation and explainability analysis of the predictive model developed using the LLM and automated ML framework
FIGURE 6 Performance evaluation and explainability analysis of the predictive model developed using the LLM and automated ML framework. (a) ROC; (b) Feature importance; (c) SHAP analysis. Abbreviations: alb_base, baseline albumin; hypoterm, intraoperative hypothermia; urea base, baseline urea; ASA, American Society of Anesthesiologists Physical Status classification; create base, baseline creatinine.

From a practical and operational standpoint, the predictive models developed and validated in this study can be seamlessly integrated into existing clinical infrastructures, such as electronic health record systems, enabling real-time risk assessment at the point of care. Integrating these predictive analytics into clinical workflows could significantly enhance postoperative management strategies, guide rehabilitation planning, and optimise resource allocation decisions. The timely identification of high-risk elderly patients undergoing hip fracture surgery could prompt targeted interventions, intensified monitoring, proactive rehabilitation strategies, and personalised follow-up plans, potentially leading to improved clinical outcomes and enhanced patient satisfaction.

Nevertheless, the findings from our study must be cautiously interpreted within the scope of several inherent limitations. First, our patient cohort was exclusively derived from a single institution, potentially limiting generalisability and introducing institution-specific biases. Second, our analysis was confined to static, routinely available clinical variables, excluding potentially informative dynamic variables such as serial vital signs, evolving laboratory markers, or longitudinal patient status data, which may enhance predictive performance when integrated. Third, our dataset represents a relatively limited timeframe, and evolving treatment paradigms, changes in clinical practice, or shifting patient demographics could influence future model performance. Consequently, periodic updating, external validation in independent patient populations, and longitudinal model recalibration are essential for ensuring continued clinical relevance and robustness. Finally, an additional consideration is that the present model was developed for all-cause one-year mortality, which represents a broad binary endpoint. If sufficiently detailed and reliable cause-of-death data were available, future studies could examine whether the prediction of specific mortality categories would improve model discrimination and clinical interpretability. It is plausible that different clinical, laboratory, and perioperative predictors may be differentially associated with distinct causes of death, and that more specific endpoint definitions could strengthen predictive performance. However, such an approach would require adequately sized cause-specific outcome groups to ensure robust model development and validation.

Taken jointly, our findings demonstrate that accurate prediction of one-year mortality in elderly patients undergoing hip fracture surgery can be effectively achieved using both traditional ML and DL methods. Among the evaluated approaches, the manually developed XGB model delivered superior predictive performance. Notably, an automated ML workflow facilitated by an LLM achieved comparable results, highlighting the potential to democratise predictive analytics for clinical practitioners without advanced technical expertise. The integration of such predictive tools into routine perioperative clinical practice represents a promising avenue for enhancing risk stratification, informing clinical decision-making, and ultimately improving patient outcomes in this vulnerable population.

Data availability statement

The raw data supporting the conclusions of this article will be made available by the authors, without undue reservation.

Ethics statement

The studies involving humans were approved by Sheba Medical Center, Israel (Approval No. SMC-D 0976-24). The studies were conducted in accordance with the local legislation and institutional requirements. The ethics committee/institutional review board waived the requirement of written informed consent for participation from the participants or the participants’ legal guardians/next of kin because due to the retrospective observational design of the study, the requirement for informed consent was waived by the committee.

Author contributions

AS: Conceptualization, Formal analysis, Investigation, Methodology, Project administration, Software, Visualization, Writing – original draft, Writing – review & editing. MG: Data curation, Formal analysis, Methodology, Validation, Writing – review & editing. MK: Investigation, Validation, Writing – review & editing. YP: Investigation, Methodology, Validation, Writing – review & editing.

HB: Resources, Supervision, Validation, Writing – review & editing. DO: Data curation, Resources, Validation, Writing – review & editing. TL: Conceptualization, Data curation, Formal analysis, Investigation, Methodology, Software, Supervision, Visualization, Writing – original draft, Writing – review & editing.

Funding

The author(s) declared that financial support was not received for this work and/or its publication.

Acknowledgments

The authors wish to thank Gal Abadi for his technical assistance with the model development.

Conflict of interest

The author(s) declared that this work was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.

Supplementary material

The Supplementary material for this article can be found online at: https://www.frontiersin.org/articles/10.3389/fmed.2026.1804645/full#supplementary-material 14.Kleinbaum DG, Dietz K, Gail M, Klein M. Logistic Regression. New York: Springer- Verlag (2002).

15.Webb GI, Keogh E, Miikkulainen R. Naïve Bayes. Encycl Mach Learn. (2010) 15:713–4.

16.Peterson LE. K-Nearest Neighbor. Scholarpedia. (2009) 4:1883. doi: 10.4249/scholarpedia.1883 17.Suthaharan S. "Support Vector Machine". In: Machine Learning Models and Algorithms for Big Data Classification. Cham: Springer (2016). p. 207–35.

18.Lazebnik T, Bunimovich-Mendrazitsky S. Decision tree post-pruning without loss of accuracy using the SAT-PP algorithm with an empirical evaluation on clinical data. Data Knowl Eng. (2023) 145:102173. doi: 10.1016/j.datak.2023.102173 19.Kaur P, Kumar R, Kumar M. A healthcare monitoring system using random forest and internet of things (IoT). Multimed Tools Appl. (2019) 78:19905–16. doi: 10.1007/s11042-019-7327-8 20.Chen T, Guestrin C. Xgboost: a scalable tree boosting system. Proceedings of the 22nd ACM SIGKDD. (2016);785–794.

21.Naraei P, Abhari A, Sadeghian A. Application of multilayer perceptron neural networks and support vector machines in classification of healthcare data. IEEE FTC. (2016) 2016:848–52.

22.Arik SÖ, Pfister T. Tabnet: attentive interpretable tabular learning. AAAI Conf Artif Intell. (2021) 35:6679–87.

23.Huang X, Khetan A, Cvitkovic M, Karnin Z. Tabtransformer: Tabular data modeling using contextual embeddings. arXiv [Preprint]; (2020). doi: 10.48550/arXiv.2012.06678 24.Alibrahim H, Ludwig SA. Hyperparameter optimization. IEEE congress. Evol Comput. (2021) 2021:1551–9.

25.Chawla NV, Bowyer KW, Hall LO, Kegelmeyer WP. SMOTE: synthetic minority oversampling technique. J Artif Intell Res. (2002) 16:321–57. doi: 10.1613/jair.953 26.Altmann A, Toloşi L, Sander O, Lengauer T. Permutation importance. Bioinformatics. (2010) 26:1340–7. doi: 10.1093/bioinformatics/btq134 27.Nohara Y, Matsumoto K, Soejima H, Nakashima N. Explanation of ML models using SHAP. Comput Methods Prog Biomed. (2022) 214:106584. doi: 10.1016/j. cmpb.2021.106584 28.Rosenfeld A, Lazebnik T. LLM Attribution for GPT Models. (2024).

29.Soffer S, Sorin V, Nadkarni G, Klang E. ChatGPT-o1 in medical ethics. medRxiv; [Preprint] (2024). doi: 10.1101/2024.09.25.24314342 30.Olson RS, Moore JH. TPOT: automating machine learning. PMLR Workshop AutoML. (2016);66–74.

31.Lazebnik T, Fleischer T, Yaniv-Rosenfeld A. Benchmarking biologically-inspired automatic machine learning for economic tasks. Sustainability. (2023) 15:11232. doi: 10.3390/su151411232 32.Fati SM, Muneer A, Akbar NA, Taib SM. A continuous cuffless blood pressure estimation using tree-based pipeline optimization tool. Symmetry. (2021) 13:686. doi: 10.3390/sym13040686 33.White Jules, Quchen Fu, Hays Sam, Sandborn Michael, Olea Carlos, Gilbert Henry, et al. A prompt pattern catalog to enhance prompt engineering with Chatgpt. arXiv [Preprint] (2023). doi: 10.48550/arXiv.2302.11382 34.Shmuel Assaf, Glickman Oren, Lazebnik Teddy. "A comprehensive benchmark of machine and deep learning across diverse tabular datasets." arXiv [Preprint] (2024). doi: 10.48550/arXiv.2408.14817 35.Xu Y, Han D, Huang T, Zhang X, Hua L, Shen S, et al. Predicting ICU mortality in rheumatic heart disease: comparison of XGBoost and logistic regression. Front Cardiovasc Med. (2022) 9:847206. doi: 10.3389/fcvm.2022. 847206 36.Aharonoff GB, Koval KJ, Skovron ML, Zuckerman JD. Hip fractures in the elderly: predictors of one year mortality. J Orthop Trauma. (1997) 11:162–5. doi: 10.1097/00005131-199704000-00004 37.Borge SJ, Lauritzen JB, Jørgensen HL. Hypoalbuminemia is associated with 30-day mortality in hip fracture patients independently of body mass index. Scand J Clin Lab Invest. (2022) 82:571–5. doi: 10.1080/00365513.2022.2150982 38.Lewis JR, Hassan SK, Wenn RT, Moran CG. Mortality and serum urea and electrolytes on admission for hip fracture patients. Injury. (2006) 37:698–704. doi: 10.1016/j. injury.2006.04.121 39.Mroczek TJ, Prodromidis AD, Pearce A, Malik RA, Charalambous CP. Perioperative hypothermia is associated with increased 30-day mortality in hip fracture patients in the United Kingdom: α systematic review and Meta-analysis. J Orthop Trauma. (2022) 36:343–8. doi: 10.1097/BOT.0000000000002332 40.Solomon A, Glebov M, Lazebnik T. Explainable surgical procedures recommender system leveraging large language models. ACM Trans Recom Syst. (2025). doi: 10.1145/3767326 41.Taskin B, Xie W, Lazebnik T. Knowledge integration for physics-informed symbolic regression using pre-trained large language models. Sci Rep. (2026) 16:1614. doi: 10.1038/s41598-026-35327-6 42.Silva Jonathan, Ma Qin, Cabot Jordi, Kelsen Pierre, Proper Henderik A. "Towards human-in-the-loop LLM-enabled domain modeling." In International Conference on Conceptual Modeling, pp. 127–145. Cham: Springer Nature Switzerland, (2025).

Article notes

Publication history
Received 5 February 2026 · Accepted 21 April 2026 · Published 1 June 2026

References

  1. Gullberg B, Johnell O, Kanis JA. World-wide projections for hip fracture. Osteoporos Int. (1997) 7:407–13. doi:10.1007/pl00004148
  2. Sanderson-Jerome C, Hariharan S. Outcome and cost evaluation of hip fractures in elderly patients at a tertiary Care Hospital in the Caribbean. Cureus. (2024) 16:e74586. doi:10.7759/cureus.74586
  3. Abrahamsen B, van Staa T, Ariely R, Olson M, Cooper C. Excess mortality following hip fracture: a systematic epidemiological review. Osteoporos Int. (2009) 20:1633–50. doi:10.1007/s00198-009-0920-3
  4. Mehta D, Gonzalez XT, Huang G, Abraham J. Machine learning-augmented interventions in perioperative care: a systematic review and meta-analysis. Br J Anaesth. (2024) 133:1159–72. doi:10.1016/j.bja.2024.08.007
  5. Schwalbe N, Wahl B. Artificial intelligence and the future of global health. Lancet. (2020) 395:1579–86. doi:10.1016/S0140-6736(20)30226-9
  6. Cary MP Jr, Zhuang F, Draelos RL, Pan W, Amarasekara S, Douthit BJ, et al. Machine learning algorithms to predict mortality and allocate palliative Care for Older Patients with hip Fracture. J Am Med Dir Assoc. (2021) 22:291–6. doi:10.1016/j.jamda.2020.09.025
  7. Li Y, Chen M, Lv H, Yin P, Zhang L, Tang P. A novel machine-learning algorithm for predicting mortality risk after hip fracture surgery. Injury. (2021) 52:1487–93. doi:10.1016/j.injury.2020.12.008
  8. Lo C-L, Yang Y-H, Hsu C-J, Chen CY, Huang WC, Tang PL, et al. Development of a mortality risk model in elderly hip fracture patients by different analytical approaches. Appl Sci. (2020) 10:6787. doi:10.3390/app10196787
  9. Riley RD, Ensor J, Snell KIE, Harrell FE, Martin GP, Reitsma JB, et al. Calculating the sample size required for developing a clinical prediction model. BMJ. (2020) 368:m441. doi:10.1136/bmj.m441
  10. Fushiki T. Estimation of prediction error by using K-fold cross-validation. Stat Comput. (2011) 21:137–46. doi:10.1007/s11222-009-9153-8
  11. Legrain A, Bouarab H, Lahrichi N. The nurse scheduling problem in real-life. J Med Syst. (2015) 39:160. doi:10.1007/s10916-014-0160-8
  12. Augustine L, Faer M, Kavountzis A, Patel R. A Brief Study of the Nurse Scheduling Problem (NSP).University of Pittsburgh Medical Center (2009).
  13. Rajeswari M, Amudhavel J, Pothula S, Dhavachelvan P. Directed bee colony optimization algorithm to solve the nurse rostering problem. Comput Intell Neurosci. (2017) 2017:6563498. Generative AI statement The author(s) declared that Generative AI was not used in the creation of this manuscript. Any alternative text (alt text) provided alongside figures in this article has been generated by Frontiers with the support of artificial intelligence and reasonable efforts have been made to ensure accuracy, including review by the authors wherever possible. If you identify any issues, please contact us. All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher. doi:10.1155/2017/6563498

This page reproduces the article Shuchami et al. (2026), Frontiers in Medicine, doi:10.3389/fmed.2026.1804645, under the CC BY licence. Text, tables and figures were extracted from the PDF and the layout adapted for the web; the PDF is the version of record.

Cite this paper

APA

Shuchami, A., Glebov, M., Katsin, M., Portnoy, Y., Berkenstadt, H., Orkin, D., & Lazebnik, T. (2026). Comparing Manual vs. Automated Machine Learning and Deep Learning Models for Predicting One-Year Mortality in Elderly Hip Fracture Patients. Frontiers in Medicine. https://doi.org/10.3389/fmed.2026.1804645

BibTeX

@article{shuchami2026comparing,
  title = {Comparing Manual vs. Automated Machine Learning and Deep Learning Models for Predicting One-Year Mortality in Elderly Hip Fracture Patients},
  author = {Shuchami, Adi and Glebov, Maxim and Katsin, Maksim and Portnoy, Yotam and Berkenstadt, Haim and Orkin, Dina and Lazebnik, Teddy},
  journal = {Frontiers in Medicine},
  year = {2026},
  doi = {10.3389/fmed.2026.1804645}
}