Assessment of Early Diabetes Risk through a Random Forest Classifier
DOI:
https://doi.org/10.56042/jsir.v85i5.27426Keywords:
Clinical decision support systems, Data mining, Diabetes prediction, Predictive analytics, Random forestAbstract
The purpose of the study was to develop an efficient machine-learning model for early detection of diabetes employing Random Forest algorithm based on a systematic workflow KDD. The findings indicate the model has good overall diagnostic utility and consistent predictive performance. The outcomes showed a reliable performance of case classification 85% accuracy, 78% specificity, 75% recall and 86% AUC for the proposal random forest model, which is good in distinguishing diabetic & nondiabetic cases. This higher specificity translates to fewer false-positive predictions, which is necessary for good clinical screening. The high AUC value is also evidence of consistent classification ability across various threshold levels. The model addressed clinical data variability successfully with an efficient machine learning solution where complex feature engineering or advanced data balance techniques were neither necessary nor segment specific, which makes it appropriate for application in real-world healthcare. The proposed framework and tools will lay the foundation for future developments, such as validation on larger, more diverse datasets, sensitivity improvements using optimization techniques and deployment through APIs for real-time healthcare systems. These results underscore the promise of structured machine learning approaches for early diabetes risk evaluation and clinical decision support.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Journal of Scientific & Industrial Research (JSIR)

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.