Assessment of Early Diabetes Risk through a Random Forest Classifier

Authors

  • Pragati Choudhari School of Computer Science Engineering & Applications, D Y Patil International University, Akurdi, Pune 412 101, Maharashtra, India
  • Shalini Goel Department of Artificial Intelligence & Data Science, IIMT College of Engineering, Greater Noida 201 310, Uttar Pradesh, India
  • Sinu Nambiar Department of Computer Engineering & Technology, Dr Vishwanath Karad MIT World Peace University, Pune 411 038, Maharashtra, India
  • Sushama Shirke Department of Computer Engineering, Army Institute of Technology, Dighi Hills, Pune 411 015, Maharashtra, India
  • Rupali Mangrule Department of Computer Science and Engineering, Maharashtra Institute of Technology, Chhatrapati Sambhajinagar 431 010, Maharashtra, India
  • Jyoti Deone Department of Information Technology, D.Y. Patil Deemed to be University RAIT, Navi Mumbai 400 706, Maharashtra, India
  • Anant Sidhappa Kurhade Department of Mechanical Engineering, Dr. D. Y. Patil Institute of Technology, Sant Tukaram Nagar, Pimpri, Pune – 411 018, Maharashtra, India

DOI:

https://doi.org/10.56042/jsir.v85i5.27426

Keywords:

Clinical decision support systems, Data mining, Diabetes prediction, Predictive analytics, Random forest

Abstract

The purpose of the study was to develop an efficient machine-learning model for early detection of diabetes employing Random Forest algorithm based on a systematic workflow KDD. The findings indicate the model has good overall diagnostic utility and consistent predictive performance. The outcomes showed a reliable performance of case classification 85% accuracy, 78% specificity, 75% recall and 86% AUC for the proposal random forest model, which is good in distinguishing diabetic & nondiabetic cases. This higher specificity translates to fewer false-positive predictions, which is necessary for good clinical screening. The high AUC value is also evidence of consistent classification ability across various threshold levels. The model addressed clinical data variability successfully with an efficient machine learning solution where complex feature engineering or advanced data balance techniques were neither necessary nor segment specific, which makes it appropriate for application in real-world healthcare. The proposed framework and tools will lay the foundation for future developments, such as validation on larger, more diverse datasets, sensitivity improvements using optimization techniques and deployment through APIs for real-time healthcare systems. These results underscore the promise of structured machine learning approaches for early diabetes risk evaluation and clinical decision support.

Downloads

Published

22.08.2026

Issue

Section

Computer Sciences, Communication and Information Technology

How to Cite

Assessment of Early Diabetes Risk through a Random Forest Classifier. (2026). Journal of Scientific & Industrial Research (JSIR), 85(5), 433-444. https://doi.org/10.56042/jsir.v85i5.27426

Similar Articles

1-10 of 212

You may also start an advanced similarity search for this article.