All projects
Machine Learning

Grand Line Message Bounty Detector — NLP Classifier

Binary text classification pipeline coupling TF-IDF vectorization with Multinomial Naive Bayes to classify spam messages with 100% precision on test data.

NLPScikit-LearnTF-IDFNaive BayesPythonCLI
spam-email-classifier.local
Spam Message & Email Classifier
The Problem

Why this matters

Unsolicited promotional and fraudulent spam messages clutter user communication and pose security risks. Building a lightweight, high-precision NLP engine is essential to eliminate false alarms where critical messages are misclassified.

Architecture

System Design

Scikit-learn unified pipeline encapsulating text preprocessing, sublinear TF-IDF vectorization (`sublinear_tf=True`), and a Multinomial Naive Bayes classifier. Evaluated on 5,572 records with stratified 80/20 train/test splitting.

Data Pipeline

Handling data accurately

  • Raw dataset parsing and validation over 5,572 records (4,825 ham, 747 spam).
  • Accent normalization and sublinear TF-IDF feature extraction inside pipeline.
  • Stratified 80/20 train/test split ensuring zero data leakage.
AI Intelligence

Machine Learning & Agents

  • Multinomial Naive Bayes probabilistic intent scoring.
  • Model serialization using `joblib` for instant real-time CLI inference.
Business Value

Measurable Outcomes & Cost Savings

Achieved 95.96% accuracy and 100.00% precision on evaluated test set (0 false positives), ensuring legitimate messages are never lost.

Stack

Technologies Used

PythonScikit-LearnPandasTF-IDFJoblibPowerShell CLI
Screenshots

System in action

Roadmap

Future Improvements

  • Expand evaluation to MIME email header datasets (Enron Spam & SpamAssassin).
  • Incorporate bi-gram and tri-gram n-gram features.
  • Build a FastAPI microservice endpoint for live web UI integration.

Ready to dive deeper?

Explore the architecture, commits, and implementation details of this project.