Comparative Analysis of Ensemble Machine Learning Algorithms for Early Prediction of Chronic Kidney Disease Using Clinical Data
DOI:
https://doi.org/10.47363/4k91bm55Keywords:
Chronic Kidney Disease (CKD), Machine Learning, Random Forest, Ensemble Learning, Medical PredictionAbstract
Background: The early identification of chronic kidney disease (CKD) is vital to minimizing the disease’s progression and improving the outcomes
of patients. Machine learning is seen as a beneficial technique for improving diagnostic proficiency since it favors automatic clinical prediction.
Methods and Materials: The present study analyzed a supervised machine learning algorithm in which a publicly available dataset with 1,000 records
from patients and seven clinical predictors was studied – the predictors being creatinine level, age, blood urea nitrogen (BUN), glomerular filtration
rate (GFR), urine production, diabetes, and hypertension. The employed methods were five other classification algorithms; Decision Tree, Random
Forest, Gradient Boosting, AdaBoost, and XGBoost, the algorithms being evaluated on the basis of accuracy, precision, recall, F1-score, ROC-AUC,
confusion matrix and importance of features. Results: Random Forest was characterized with the best performance (97.3% accuracy and 0.99 ROC-AUC) while XGBoost had nearly similar results (97.0% accuracy). It was found out while looking at the importance of features that GFR, creatinine level, and BUN are the main predictors of CKD.
Discussion: Ensemble learning techniques outperformed the conventional Decision Tree method in terms of predictive performance. Given these
facts, Random Forest seems to provide the best balance between accuracy, efficiency, and interpretability, making it a good candidate for a clinical
decision-support tool.
Conclusion: Random Forest had the best successful prediction performance among all techniques used being highly effective and interpretable for CKD detection.