Generic selectors
Exact matches only
Search in title
Search in content
Post Type Selectors
Search in posts
Search in pages
Filter by Categories
Case Report
Clinical research study
Current Issue
Editorial Board
Literature Review
Narrative review
Original Article
Research Article
Review Article
Short Report
Surgical techniques
Generic selectors
Exact matches only
Search in title
Search in content
Post Type Selectors
Search in posts
Search in pages
Filter by Categories
Case Report
Clinical research study
Current Issue
Editorial Board
Literature Review
Narrative review
Original Article
Research Article
Review Article
Short Report
Surgical techniques
View/Download PDF

Translate this page into:

71 (); 291-298
doi:
10.1016/j.jor.2025.10.008

Developing and validating machine learning models to predict length of hospitalization before obese patients undergo elective arthroplasty

Hospital for Special Surgery, New York, NY, USA
Department of Orthopedic Surgery, Balgrist University Hospital, University of Zürich, Zurich, Switzerland

⁎Corresponding author: Felix C. Oettl. felix.oettl@balgrist.ch

Disclaimer:
This article was originally published by Reed Elsevier India Pvt. Ltd. and was migrated to Scientific Scholar after the change of Publisher.

Abstract

Abstract

Obese patients undergoing Total Joint Arthroplasty (TJA) have been associated with increased length of hospital stay (LOS) and in-hospital resource utilization. This poses challenges for institutions participating in bundled-payment programs. We investigated whether pre-operative information could predict prolonged hospital stays in obese patients undergoing TJA.

Using the arthroplasty registry of a single large institution, we identified 4563 obese patients who underwent unilateral THA or TKA between 2020 and 2021. No data was missing or incomplete. A total of 31 pre-surgical parameters were included as potential predictor variables. The data was partitioned into training (80 %) and test (20 %) set. For binary modelling, patients were categorized by a LOS less than 2 nights (41.4 %) and a LOS of 2 nights or more (58.6 %). Binary classification model performance was evaluated using area under the receiver operating characteristic curve (AUC-ROC), Matthews correlation coefficient (MCC), F1 score, accuracy, sensitivity, precision, and the Brier score. Regression models were evaluated utilizing Root Mean Squared Error (RMSE) and Mean Absolute Error (MAE). 95 % confidence intervals (95 % CI) were calculated via bootstrapping.

The explainable boosted machine (EBM) model was the best performing binary model: F1-Score 0.75 (95 % CI 0.72, 0.78); accuracy 0.69 (0.67, 0.72); sensitivity 0.71 (0.67, 0.75); precision 0.74 (0.7, 0.78), AUC-ROC 0.75 (0.72, 0.79), MCC 0.38 (0.32, 0.44) and Brier Score 0.199. The QRF was the best performing regression model with MAE of 21.5 (20.0, 23.2) and RMSE of 32.3 (28.8, 40.0) hours. Male sex, older age, and not being married were the most important predictors across both models. Correlation of QRF predicted and actual LOS revealed a Spearman Correlation Coefficient of 0.43 (0.36, 0.47).

All models demonstrated strong predictive capability, underscoring their clinical relevance. These insights can guide preoperative planning, patient counseling, and resource allocation, helping optimize care and discharge strategies.

1

1 Introduction

Total Joint Arthroplasty (TJA), which includes Total Hip and Total Knee Arthroplasty (THA and TKA), stands as one of the most successful surgical interventions in modern medicine. These procedures have profoundly transformed the lives of millions of patients, alleviating the debilitating effects of osteoarthritis (OA) and restoring quality of life.1,2 As the population continues to age and the rate of obesity rises, the demand for these life-changing procedures has surged, underscoring the need to address the mounting burden of musculoskeletal disorders.3

Healthcare policies have encouraged ambulatory surgery, aiming for shorter hospital stays and fewer complications. Identifying predictive variables of length of stay (LOS) is crucial for delivering quality and cost-effective arthroplasty care in this evolving landscape.4,5 Obesity, a well-established risk factor for a variety of poor health outcomes, has been inextricably linked to increased postoperative complications and prolonged LOS following joint replacement surgery.6–10

Machine learning (ML) techniques have demonstrated promising potential in predicting LOS and identifying risk factors for prolonged hospitalization across various clinical settings.4,11 While previous studies have included obesity as a variable in predictive models for LOS, to our knowledge, this is the first study to focus specifically on developing and validating machine learning models for a cohort composed exclusively of obese patients. This is particularly relevant in high-volume centers where non-obese patients are increasingly treated on an ambulatory basis, creating a distinct clinical population that would confound a combined analysis.

This study aims to construct and validate models that can accurately predict the duration of hospitalization for obese patients undergoing elective THA and TKA. Providing evidence-based guidance on the expected LOS can inform perioperative decision-making, resource allocation, and risk stratification, ultimately optimizing the delivery of cost-effective, high-quality arthroplasty care for this vulnerable patient population.

2

2 Material and methods

2.1

2.1 Design

This study was conducted at a single academic center. The study adhered to the Transparent Reporting of a Multivariable Prediction Model for Individual Prognosis or Diagnosis (TRIPOD) guidelines as well as the Guidelines for Developing and Reporting ML Models in Biomedical Research.12,13 The collection of clinical data was undertaken with the approval of the institutional review board and ethics committee at our institution (IRB #2022-1668), ensuring compliance with ethical standards and protocols.

The study was performed using the arthroplasty registry at our institution, comprised of all patients undergoing joint arthroplasty at our institution. Inclusion criteria were (1) primary THA or TKA (2) BMI ≥30 (3) TJA between January 2020 and December 2021. After applying our inclusion criteria 4563 obese patients who underwent unilateral THA or TKA between 2020 and 2021 were identified and included in the analysis. The study period of 2020–2021 was chosen to minimize the influence of historical changes in surgical techniques, pain management protocols, and discharge criteria, thus providing a model that is more relevant to current clinical practice.

The outcome of interest is patient length of stay, both as a dichotomous outcome defined as 2 days or more and as a continuous outcome in hours. The binary cutoff of a length of stay of 2 nights or more was determined based on the consensus of senior surgeons at our institution, who consider this a clinically meaningful threshold for a prolonged stay in this patient population.

2.2

2.2 Source data and input

All prediction models were developed using 31 categorical or continuous variables (supplementary material), all of which are demographic or planned pre-procedure and therefore available before the patient is admitted. For the models unable to receive categorical variables as input, variables were one hot encoded before training the respective models. The analysis of variable importances showed the extent to which each variable affects the model's ability to make predictions. The top 3 variables exhibiting the highest degree of importance across both classification and regression models were identified as key variables driving the performance of the models in this study. No data was missing or incomplete. The dataset was randomly split into training (80 %) and test (20 %) set, with a set seed of 42 to ensure reproducibility.

2.3

2.3 Model development

Categorical variables were summarized by presenting their absolute frequencies and corresponding percentages, whereas continuous variables were summarized means and standard deviations, as depicted in Table 1. The predictive modeling for the dichotomous model was performed by employing three distinct ML algorithms: explainable boosted machine (EBM),14 extreme gradient boosting (XGBoost), and logistic regression (LR). For the continuous outcome model an EBM, XGBoost, and Quantile Regression Forests (QRF)15,16 were utilized. EBM and XGBoost are ensemble methods based on decision trees, with the latter gaining popularity in recent years due to high accuracy on complex datasets. QRF is a non-parametric ensemble methodology that utilizes tree-based models to estimate conditional quantiles.15 It is a generalization of the random forests algorithm, a versatile ensemble learning algorithm, which has gained widespread popularity and proven highly useful as a general-purpose ML method.

Table 1 Characteristics of the Training and Testing Sets and complete cohort.
Characteristic Complete cohort (N = 4563) Training Set (n = 3650) Testing Set (n = 913)
Length of stay 57.34 ± 39.37 57.73 ± 40.48 55.75 ± 34.53
2 or more nights n (%) 2675 (58.6) 2154 (59) 521 (57.1)
Age 63.97 ± 9.1 64.04 ± 9.02 63.72 ± 9.41
BMI 35.2 ± 4.2 35.3 ± 4.3 35.1 ± 4.2
Male Sex, n (%) 1997 (43.8) 1592 (43.6) 405 (44.4)
Surgery year, n (%)
2020 2563 (58.13) 2117 (58) 536 (58.7)
2021 1910 (41.85) 1533 (42) 377 (41.3
First Race, n (%)
White or Caucasian 3645 (79.9) 2904 (79.6) 741 (81.2)
Black or African American 473 (10.4) 399 (10.9) 74 (8.1)
Asian 72 (1.6) 55 (1.5) 17 (1.9)
American Indian or Alaska native 11 (0.2) 6 (0.2) 5 (0.5)
Native Hawaiian or Other Pacific Islander 1 (<0.1) 0 1 (0.1)
Other 275 (6) 223 (6.1) 52 (5.7)
Unavailable 86 (1.9) 63 (1.7) 23 (2.5)
Marital Status, n (%)
Married 3036 (66.5) 2412 (66.1) 624 (68.3)
Single 689 (15.1) 567 (15.5) 122 (13.4)
Divorced 368 (8.1) 298 (8.2) 70 (7.7)
Widowed 357 (7.8) 282 (7.7) 75 (8.2)
Significant other 48 (1.1) 44 (1.2) 4 (0.4)
Legally separated 27 (0.6) 18 (0.5) 9 (1)
Other 48 (1.1) 10 (0.3) 7 (0.8)
Unknown/Unspecified 21 (0.5) 19 (0.5) 2 (0.2)
Joint, n (%)
Knee 2641 (57.9) 2128 (58.3) 513 (56.2)
Hip 1922 (42.1) 1522 (41.7) 400 (43.8)
ASA Level, n (%)
I 34 (0.7) 28 (0.7) 6 (0.7)
II 3209 (70.3) 2570 (70.4) 6,39 (70)
III+ 1320 (29) 1052 (28.9) 268(30.2)
Charlson comorbidity index, n (%)
CCI 0 2826 (61.9) 2256 (61.8) 570 (62.4)
CCI 1 1136 (24.9) 923 (25.3) 213 (23.3)
CCI 2 363 (7.9) 278 (7.6) 85 (9.3)
CCI 3 121 (2.7) 100 (2.7) 21 (2.3)
CCI 4 69 (1.5) 54 (1.5) 15 (1.6)
CCI 5+ 48 (1.1) 39 (1.1) 8 (1)
Smoking status, n (%)
Current Smoker 247 (5.4) 202 (5.5) 45 (4.9)
Former Smoker 1651 (36.2) 1334 (36.6) 317 (34.7)
Never Smoker 2655 (58.2) 2105 (57.7) 550 (60.3)
Current Status Unknown 10 (0.2) 9 (0.2) 1 (0.1)
State, n (%)
New York 3528 (77.3) 2829 (77.5) 699 (76.6)
New Jersey 880 (19.3) 699 (19.2) 181 (19.8)
Connecticut 106 (2.3) 83 (2.2) 23 (2.5)
Pennsylvania 49 (1.1) 39 (1.1) 10 (1.1)
Hip approach, n (%)
Posterior 1687 (37) 1338 (36.7) 349 (38.2)
Anterior 235 (5.2) 184 (5) 51 (5.6)
Not applicable (TKA) 2641 (57.8) 2128 (58.3) 513 (56.2)

To prevent overfitting, we employed 7-fold cross-validation during model training. Hyperparameters for each model were tuned using a grid search algorithm to identify the optimal combination for predictive performance.

To evaluate a representative diversity of methods with differing characteristics, we selected these 3 models for the respective predictive approach. The 6 algorithms were derived using the training cohort and validated on the test cohort.

The machine learning models used in this study were selected based on their established performance in clinical prediction tasks and, in the case of the Explainable Boosting Machine and logistic regression, their inherent interpretability. While more complex models such as neural networks were considered, we prioritized models that would allow for a clear understanding of the factors driving the predictions.

2.4

2.4 Model selection

Predictive performance and model interpretability were assessed to determine the best model. For classification models Matthews Correlation Coefficient (MCC), Area under the receiver operating curve (AUC-ROC), F1-Score (harmonic mean of precision and recall), Accuracy, Sensitivity, Precision, and Brier score were calculated as performance metrics. Regression model performance was assessed with Root Mean Squared Error (RMSE) and Mean Absolute Error (MAE). 95 % confidence intervals (95 % CI) were calculated via bootstrapping.

We performed the analysis and model construction utilizing Python packages (v3.11, lightgbm, matplotlib, numpy, pandas, sklearn, xgboost, InterpretML, RandomForestQuantileRegressor).

3

3 Feature importance

Feature weights are calculated from the models, which quantify the relative importance of each variable in predicting the outcome variable. Additionally, partial dependency plots were generated for the weighted features of the EBM model to examine the functional relationship between the predictors and the outcome at the population level. Shapley values were calculated for the QRF. These plots facilitate the investigation of the directionality and shape of the associations, providing valuable insights for patients, surgeons, and other stakeholders involved in the clinical decision-making process.

4

4 Results

The study encompassed a total cohort of 4563 patients, 58.9 % of which were hospitalized for a duration of two days or more. This patient population was then randomly split into two groups: a test set comprising 913 patients and a training set of 3650 patients. The distribution of variables in these groups is illustrated in Table 1.

4.1

4.1 Prediction model

For the classification task, the EBM, XGBoost, and LR models exhibited comparable performance without a distinct advantage for any particular model upon assessment. However, we elected to report more in-depth findings on the EBM, as it is an inherently transparent and explainable model (Fig. 1,Table 2).

AUC-ROC of the EBM model.
Fig. 1 AUC-ROC of the EBM model.
Table 2 Performance of Classification Models on test Set.
Model AUC-ROC (95 %CI) MCC (95 %CI) F1-Score (95 %CI) Accuracy (95 %CI) Sensitivity (95 %CI) Precision (95 %CI) Brier - score
EBM 0.75 (0.72, 0.78) 0.37 (0.31, 0.44) 0.72 (0.69, 0.75) 0.69 (0.66, 0.72) 0.71 (0.67, 0.75) 0.74 (0.7, 0.78) 0.1993
XGBoost 0.75 (0.72, 0.79) 0.38 (0.31, 0.44) 0.73 (0.7, 0.76) 0.69 (0.67, 0.72) 0.74 (0.7, 0.77) 0.73 (0.69, 0.77) 0.2013
Logistic reg. 0.75 (0.72, 0.78) 0.38 (0.32, 0.45) 0.72 (0.68, 0.75) 0.69 (0.66, 0.72) 0.68 (0.64, 0.72) 0.75 (0.71, 0.75) 0.1982

For the regression task of predicting the precise length of stay down to the hour, all models exhibited satisfactory performance. However, the QRF model demonstrated superior predictive ability as it was the sole model that captured the cyclic patterns inherent to patient admissions and discharges (Table 3).

Table 3 Performance of Regression Models on test Set.
Model MAE (95 %CI) measured in hours RMSE (95 %CI) measured in hours Spearman correlation coefficient (95 % CI)
EBM 21.9 (20.68, 23.41) 30.77 (27.60, 35.31) 0.48 (0.43, 0.53)
Elastic Net 21.86 (20.58, 23.38) 30.56 (27.34, 34.82) 0.48 (0.43, 0.53)
XGBoost 22.31 (20.99, 23.82) 31.4 (28.11, 35.86) 0.47 (0.42, 0.52)
QRF 21.26 (19.84, 22.96) 32.23 (28.82, 36.86) 0.43 (0.36, 0.47)
4.2

4.2 Feature importance

The features demonstrating the greatest influence on the EBM model were sex, age, marital status, BMI, and race (Fig. 2). Analysis of partial dependency plots revealed that an increased probability of a hospital stay lasting two or more days was associated with female sex, advanced age, non-married marital status, higher BMI, and non-white race. Patients older than 69.5 years and those with a BMI greater than 40 kg/m2 exhibited an increased likelihood of prolonged admission. Interestingly, the partial dependency plot of BMI did not demonstrate a linear relationship, but rather distinct peaks and valleys at BMI values of 31, 33, and 40 (Fig. 3).

Features ranked by importance for prediction (mean absolute scores).
Fig. 2 Features ranked by importance for prediction (mean absolute scores).
a = Partial dependency plot of Age; b = Partial dependency plot of BMI at surgery.
Fig. 3 a = Partial dependency plot of Age; b = Partial dependency plot of BMI at surgery.

In the QRF model, sex, age, BMI, type of procedure, and race were the most important predictors. The model structure does not allow for partial dependency plots.

4.3

4.3 Correlation of regression prediction

Upon completing the regression analysis to predict each patient's length of stay, the model outputs were correlated with the actual length of stay using the Spearman correlation coefficient. The QRF did show inferior correlation coefficient compared to the other model; however, the QRF was the sole model capable of capturing the cyclical patterns inherent to patient admissions and discharges (Fig. 4).

Scatterplot comparing predicted and actual length of stay for the 4 regression models. a = explainable boosted machine; b = elastic net; c = XGBoost d = Quantile Regression Forest.
Fig. 4 Scatterplot comparing predicted and actual length of stay for the 4 regression models. a = explainable boosted machine; b = elastic net; c = XGBoost d = Quantile Regression Forest.
5

5 Discussion

In this study, we developed a ML model to predict LOS following primary TJA. We evaluated 7 ML algorithms: 3 for classification and 4 for regression.

This study has several limitations that should be considered when interpreting the results. As a single-center retrospective study focused solely on obese patients undergoing primary total joint arthroplasty, the generalizability of the findings may be limited to similar healthcare settings and patient populations. The retrospective design is subject to potential biases, and the limited feature set of 31 variables may not capture all relevant factors influencing length of stay. We were not able to include perioperative, socioeconomic, or other institutional factors that may also influence length of stay. Future studies should aim to incorporate a wider range of variables to improve the predictive performance of the models. The lack of external validation on an independent dataset raises concerns about the true performance and generalizability of the models. While following the TRIPOD guidelines, the single-center design, limited patient population, lack of external validation, and potential biases should be considered when interpreting and applying the results.12

All tested classification models showed comparable performance, thus we decided to further examine the EBM model due to its superior explainability characteristics, natively handling missingness and a good AUC-ROC of 0.75 (Fig. 1, Table 2), indicating that the model accurately discriminates between patients who did and did not stay 2 or more nights. The discrepancies between the models were more pronounced in the regression models, where the respective models showed varied performances dependent on the metric tested. Due to the left-skewed nature of our data with extreme outliers, the MAE––which is less sensitive to outliers compared to the RMSE––was used to determine which model would be further examined. Based on this assumption, the QRF slightly outperformed the competing models, however with overlapping confidence intervals. We found that sex, age, marital status, BMI, and race were the most important predictors for LOS (Fig. 2). Moreover, the associated partial dependency plot of age and BMI showed nonlinear associations with the probability of prolonged LOS, making it possible to potentially establish cutoff points facilitating decision-making. To our knowledge, this is the largest single center study using ML to predict Length of Stay in the obese population.

Previous studies on LOS prediction explored the association of risk factors with prolonged length of stay.11,17,18 Our findings align closely with established risk factors in the literature, reinforcing that sex, age, marital status, BMI, and race are key predictors of length of stay. These factors, however, cannot be easily altered at the time of surgery.

Our work found that male sex, being married, and White or Caucasian race are significant predictors for patients staying less than 2 nights (Figs. 5 and 6 a). Regarding age, we see a non-linear relationship with risk of prolonged admission, with a constant risk up to 70 years, after which we see an almost linear increase in risk for prolonged admission until 84 years, where the risk curve flattens out (Fig. 3). Conversely, with BMI we see 2 steps with increased risk for prolonged LOS at a BMI of 40.15 and 47.25 respectively (Fig. 3).

a = Partial dependency plot of Sex; b = Partial dependency plot of Martial status.
Fig. 5 a = Partial dependency plot of Sex; b = Partial dependency plot of Martial status.
a = Partial dependency plot of Race; b = Partial dependency plot of Procedure.
Fig. 6 a = Partial dependency plot of Race; b = Partial dependency plot of Procedure.

The strong fluctuations around the BMI values of 31, 33, and 40 are likely related to changes in protocol of high-volume surgeons (Fig. 3). While only the 6th most important variable in the EBM model, procedure type shows a clear trend between THA and TKA, as well as computer assisted vs manual procedures, with computer-assisted THA showing the lowest risk of a prolonged LOS and manual TKA the highest (Fig. 6 a). Our findings could be used to inform risk adjustment in future studies comparing LOS as they further support existing literature.

Notably, our feature importance analysis displayed Year of Surgery as the 5th most important variable, showing a shorter length of stay in patients treated in 2021 compared to 2020 (Fig. 2). The decrease in LOS has been evident since 2016, prompting us to exclude data prior to 2020 to ensure a more accurate assessment of risk factors. This exclusion prevents the LOS prediction from being unduly influenced by changes in healthcare policy. Additional research in the general population undergoing TJA is warranted to report a more precise prediction in the non-obese population and assess if the reported risk factors are robust in other populations.

It is important to acknowledge the ethical considerations associated with using non-modifiable predictors such as sex and race. These variables should not be used to restrict access to care or to reinforce existing health disparities. Instead, they should be used to identify patients who may be at higher risk for a prolonged length of stay and who may benefit from more intensive preoperative optimization, personalized patient education, and more comprehensive discharge planning. The goal of these models is to facilitate more equitable and efficient allocation of healthcare resources, not to create new barriers to care.

Among the continuous models evaluated, the QRF model demonstrated notable performance in predicting patient length of stay following total joint arthroplasty. While the QRF did not outperform other models in terms of mean average error (MAE), it exhibited a unique capability to capture the cyclical patterns seen in patient admissions and discharges (Table 3, Fig. 4). The correct prediction of cyclical patterns in length of stay by the QRF model is an intriguing finding, as the model is less likely to predict a discharge in the middle of the night. The ability of the QRF model to capture these patterns comes from its ability to estimate conditional quantiles and highlights the potential value of non-parametric, ensemble-based approaches in healthcare predictive modeling. The significance of accurately capturing a specific pattern in hospital discharge hours (Fig. 4), allows for a more accurate planning, as compared to a model without this ability technical performance might be higher, however with reduced real world applicability. While traditional regression models may struggle to identify complex, non-linear relationships, the QRF's tree-based ensemble structure allows it to adapt to and uncover intricate patterns within the data. The ability to predict LOS before surgery offers several actionable benefits for hospital systems. For patients identified as having a high likelihood of prolonged hospitalization, targeted interventions such as optimizing modifiable preoperative factors could be considered to reduce risk. Additionally, hospitals can use these predictions for capacity planning, ensuring adequate bed availability and resource allocation to improve patient flow and reduce bottlenecks in high-volume centers. In cases where prolonged hospitalization is anticipated and hospital resources are constrained, alternative treatment facilities, such as specialized rehabilitation centers or step-down units, could be considered to better accommodate patient needs.

All models demonstrated moderate predictive capability. While the performance metrics are not perfect, they suggest that these models can be a useful adjunct to clinical judgment in preoperative planning, patient counseling, and resource allocation.

Future validation efforts should focus on assessing model performance across multi-center datasets to enhance generalizability. Testing in different hospital systems would help evaluate robustness across diverse patient populations, surgical protocols, and institutional workflows. Expanding the feature set to include additional clinical and socioeconomic factors may further improve predictive accuracy. Prospective validation, where the model is applied in real-time clinical settings to predict LOS before surgery, along with external benchmarking, will be essential to confirm real-world applicability and mitigate potential biases inherent in the retrospective design.

6

6 Conclusion

These insights can help to inform preoperative planning, patient counseling, and resource allocation. While not a substitute for clinical judgment, these models can be a valuable tool for identifying high-risk patients and for optimizing care and discharge strategies. Further research, including external validation in multi-center cohorts, is needed to confirm the generalizability of our findings and to further refine these predictive models.Legend of abbreviationsAUC-ROCarea under the receiver operating characteristic curveEBMexplainable boosted machineLOSlength of hospital stayLRlogistic regressionMAEMean Absolute ErrorMCCMatthews correlation coefficientMLMachine learningOAosteoarthritisQRFquantile regression forestRMSERoot Mean Squared ErrorTHATotal Hip ArthroplastyTJATotal Joint ArthroplastyTKATotal Knee ArthroplastyTRIPODTransparent Reporting of a Multivariable Prediction Model for Individual Prognosis or DiagnosisXGBoostextreme gradient boosting

Ethics approval and consent to participate

The collection of clinical data was undertaken with the approval of the institutional review board and ethics committee at our institution (IRB #2022-1668), ensuring compliance with ethical standards and protocols.

Consent for publication

The collection of clinical data was undertaken with the approval of the institutional review board and ethics committee at our institution (IRB #2022-1668), ensuring compliance with ethical standards and protocols.

Availability of data and material

This study utilized patient data with numerous variables. Due to the sensitive nature of this information and to ensure compliance with HIPAA regulations, any requests for data sharing will need to undergo review by our legal team before approval can be granted.

Ethical approval

N/A.

Author's contributions

All listed authors have contributed substantially to this work: FCO developed the machine learning model. FCO, AIW and JGN performed primary manuscript preparation. Editing and final manuscript preparation was performed by MP, FC, AGDV, AIW and FCO. All authors read and approved the final manuscript.

Ethical statement

This study was conducted at a single academic center and adhered to the highest ethical standards. The collection and use of clinical data were approved by the institutional review board and ethics committee at our institution (IRB #2022-1668). The study was performed in accordance with the principles of the Declaration of Helsinki and its later amendments. In line with best practices for transparency and rigor in predictive modeling research, the study followed the Transparent Reporting of a Multivariable Prediction Model for Individual Prognosis or Diagnosis (TRIPOD) guidelines and the Guidelines for Developing and Reporting Machine Learning Models in Biomedical Research. Given the retrospective nature of the study and the use of de-identified data, the institutional review board waived the requirement for individual patient consent.

Funding

None.

References

  1. , , . Burden of major musculoskeletal conditions. Bull World Health Organ. 2003;81(9):646-656.
    [Google Scholar]
  2. , , . What is the prevalence of musculoskeletal problems in the elderly population in developed countries? A systematic critical literature review. Chiropr Man Ther. 2012;20(1):31.
    [Google Scholar]
  3. American joint replacement registry (AJRR) 2023
    [Google Scholar]
  4. , , , , . Length of stay after joint arthroplasty is less than predicted using two risk calculators. J Arthroplast. 2021;36(9):3073-3077.
    [Google Scholar]
  5. , , , , , , . Why still in hospital after fast-track hip and knee arthroplasty? Acta Orthop. 2011;82(6):679-684.
    [Google Scholar]
  6. , , . Obesity, preoperative weight loss, and telemedicine before total joint arthroplasty: a review. Arthroplasty. 2022;4(1)
    [Google Scholar]
  7. , , , et al . Total hip arthroplasty outcomes in morbidly obese patients. EFORT Open Rev. 2018;3(9):507-512.
    [Google Scholar]
  8. , , , , , , . Functional gain and pain relief after total joint replacement according to obesity status. J Bone Joint Surg. 2017;99(14):1183-1189.
    [Google Scholar]
  9. , , , et al . The outcomes of total knee arthroplasty in morbidly obese patients: a systematic review of the literature. Arch Orthop Trauma Surg. 2019;139(4):553-560.
    [Google Scholar]
  10. , , , , . Greater risks of complications, infections, and revisions in the obese versus non-obese total hip arthroplasty population of 2,190,824 patients: a meta-analysis and systematic review. Osteoarthr Cartil. 2020;28(1):31-44.
    [Google Scholar]
  11. , , , , . Preoperative prediction and risk factor identification of hospital length of stay for total joint arthroplasty patients using machine learning. Arthroplasty Today. 2023;22
    [Google Scholar]
  12. , , , , . Transparent reporting of a multivariable prediction model for individual prognosis or diagnosis (TRIPOD): the TRIPOD statement. Br J Surg. 2015;102(3):148-158.
    [Google Scholar]
  13. , , , et al . Guidelines for developing and reporting machine learning predictive models in biomedical research: a multidisciplinary view. J Med Internet Res. 2016;18(12)
    [Google Scholar]
  14. , , , , . Interpretml: A Unified Framework for Machine Learning Interpretability. 2019
    [Google Scholar]
  15. , . Quantile regression forests. J Mach Learn Res. 2006;7:983-999.
    [Google Scholar]
  16. , . quantile-forest: a python package for quantile regression forests. J Open Source Softw. 2024;9(93):5976.
    [Google Scholar]
  17. , , , , , , . Construction and comparison of predictive models for length of stay after total knee arthroplasty: regression model and machine learning analysis based on 1,826 cases in a single Singapore center. J Knee Surg. 2022;35(1):7-14.
    [Google Scholar]
  18. , , , et al . Hospital length of stay prediction for general surgery and total knee arthroplasty admissions: systematic review and meta-analysis of published prediction models. Digital Health. 2023;9
    [Google Scholar]
Show Sections