Grand Line Message Bounty Detector — NLP Classifier
Binary text classification pipeline coupling TF-IDF vectorization with Multinomial Naive Bayes to classify spam messages with 100% precision on test data.

Why this matters
Unsolicited promotional and fraudulent spam messages clutter user communication and pose security risks. Building a lightweight, high-precision NLP engine is essential to eliminate false alarms where critical messages are misclassified.
System Design
Scikit-learn unified pipeline encapsulating text preprocessing, sublinear TF-IDF vectorization (`sublinear_tf=True`), and a Multinomial Naive Bayes classifier. Evaluated on 5,572 records with stratified 80/20 train/test splitting.
Handling data accurately
- Raw dataset parsing and validation over 5,572 records (4,825 ham, 747 spam).
- Accent normalization and sublinear TF-IDF feature extraction inside pipeline.
- Stratified 80/20 train/test split ensuring zero data leakage.
Machine Learning & Agents
- Multinomial Naive Bayes probabilistic intent scoring.
- Model serialization using `joblib` for instant real-time CLI inference.
Measurable Outcomes & Cost Savings
Achieved 95.96% accuracy and 100.00% precision on evaluated test set (0 false positives), ensuring legitimate messages are never lost.
Technologies Used
System in action
Future Improvements
- Expand evaluation to MIME email header datasets (Enron Spam & SpamAssassin).
- Incorporate bi-gram and tri-gram n-gram features.
- Build a FastAPI microservice endpoint for live web UI integration.
Ready to dive deeper?
Explore the architecture, commits, and implementation details of this project.