The Early Diabetes Risk Prediction Using Machine Learning system is a comprehensive machine learning solution that predicts diabetes risk using the Pima Indian Diabetes dataset. This project is ideal for B.Tech, MCA, CSE, and engineering students as their final year project. It implements Random Forest, XGBoost, and Logistic Regression with leak-free SMOTE, achieving 84.8% accuracy and 0.920 ROC-AUC.
The system utilizes the Pima Indian Diabetes dataset containing 768 patient records with 8 clinical features including glucose, BMI, age, pregnancies, insulin, blood pressure, skin thickness, and diabetes pedigree function. It provides a user-friendly web interface for data upload, exploratory data analysis, model training, and real-time diabetes risk prediction, enabling healthcare professionals to make informed clinical decisions and reduce diabetes-related complications.
| Metric | Logistic Regression | Random Forest | XGBoost | Best |
|---|---|---|---|---|
| Test Accuracy | 78.9% | 84.8% | 84.5% | Random Forest |
| Test Precision | 78.5% | 83.9% | 83.5% | Random Forest |
| Test Recall | 78.9% | 84.8% | 84.5% | Random Forest |
| Test F1-Score | 78.7% | 84.3% | 83.9% | Random Forest |
| Test ROC-AUC | 83.4% | 91.4% | 92.0% | XGBoost |
| CV Mean ROC-AUC | 82.5% | 90.6% | 91.0% | XGBoost |
Complete thesis writing, research guidance, and formatting support
Expert HelpQuality assignment writing, editing, and proofreading services
100% OriginalResearch proposal, literature review, data analysis & publication
PhD LevelAcademic projects, mini projects, and final year project support
Hands-on