This project, developed for ProfessionAI, focuses on analyzing and classifying incoming emails, with a particular emphasis on identifying SPAM messages and studying their content.
The workflow is structured into four main tasks:
-
SPAM/HAM Classification with Naive Bayes
A Naive Bayes classifier was trained to distinguish between legitimate (HAM) and SPAM emails. -
Topic Modeling on SPAM Emails
Topics were extracted from SPAM emails to identify recurring themes and patterns. -
Semantic Distance between Topics
The semantic distance between SPAM topics was calculated to evaluate the diversity of unwanted content. -
Organization Extraction from HAM Emails
Organizations mentioned in legitimate emails were identified, providing useful insights for business intelligence.
Added Value:
This pipeline improves anti-spam filters, enables a deeper understanding of SPAM trends and content, enhances communication security, and enriches decision-making processes with insights extracted from legitimate emails.
Stefano Trovato
GitHub Profile
LinkedIn
