The Toxic and Abusive Online Content Detection Using Machine Learning system is a comprehensive machine learning solution that detects toxic and abusive content in online comments. This project is ideal for B.Tech, MCA, CSE, and engineering students as their final year project. It implements Random Forest and XGBoost with leak-free SMOTE, achieving 93.8% accuracy and 86.56% F1-score.
The system utilizes the Toxic Comment Classification Dataset from Wikipedia talk pages, containing approximately 159,571 comments labeled across six toxicity categories. It provides a user-friendly web interface for data upload, exploratory data analysis, model training, and real-time toxic content prediction, enabling social media platforms and online communities to maintain safe digital environments.
| Metric | Random Forest | XGBoost | Best |
|---|---|---|---|
| Test Accuracy | 92.5% | 93.8% | XGBoost |
| Test Precision | 85.67% | 87.19% | XGBoost |
| Test Recall | 84.32% | 85.93% | XGBoost |
| Test F1-Score | 84.99% | 86.56% | XGBoost |
| Test ROC-AUC | 92.81% | 94.21% | XGBoost |
| CV Mean ROC-AUC | 92.14% | 93.58% | XGBoost |
Complete thesis writing, research guidance, and formatting support
Expert HelpQuality assignment writing, editing, and proofreading services
100% OriginalResearch proposal, literature review, data analysis & publication
PhD LevelAcademic projects, mini projects, and final year project support
Hands-on