The Hate Speech and Offensive Language Detection Using Machine Learning system is a comprehensive machine learning solution that detects and classifies tweets into Hate Speech, Offensive Language, and Neither categories. This project is ideal for B.Tech, MCA, CSE, and engineering students as their final year project. It implements XGBoost and Random Forest with enhanced feature engineering and leak-free SMOTE, achieving 95.7% accuracy and 0.951 F1-score.
The system utilizes a labeled tweet dataset with approximately 25,000 records and extracts 23 numeric features including sentiment indicators, offensive word counts, and linguistic patterns, combined with TF-IDF vectorization. It provides a user-friendly Flask-based web interface for real-time tweet analysis, batch prediction, and comprehensive reporting, offering practical applications for social media moderation, content filtering, and online safety enforcement.
| Metric | Random Forest | XGBoost | Best |
|---|---|---|---|
| Test Accuracy | 94.1% | 95.7% | XGBoost |
| Test Precision | 93.8% | 95.4% | XGBoost |
| Test Recall | 94.0% | 95.6% | XGBoost |
| Test F1-Score | 93.6% | 95.1% | XGBoost |
| Test ROC-AUC | 95.4% | 96.8% | XGBoost |
| CV Mean F1-Macro | 92.8% | 94.5% | XGBoost |
Complete thesis writing, research guidance, and formatting support
Expert HelpQuality assignment writing, editing, and proofreading services
100% OriginalResearch proposal, literature review, data analysis & publication
PhD LevelAcademic projects, mini projects, and final year project support
Hands-on