Chinta, Pavan Kumar (2025) Semantic Text Embedding Framework for Phishing Email Detection: Leveraging Pretrained Language Models for Content-Based Threat Classification. Masters thesis, Dublin, National College of Ireland.
Preview |
PDF (Master of Science)
Download (1MB) | Preview |
Preview |
PDF (Configuration Manual)
Download (1MB) | Preview |
Abstract
Phishing emails are a constant danger to cybersecurity, and AI-generated assaults can evade detection systems by more than 96%. This study compares base pre-trained transformer embeddings to fine-tuned phrase transformers and conventional feature extraction techniques in order to address the dearth of comparative evaluation of text embedding methodologies for phishing detection. A three-tier embedding framework was implemented, extracting nine distinct representations from 83,084 curated phishing and legitimate emails: traditional methods (TF-IDF, bag-of-words), base semantic embeddings (BERT, RoBERTa, DistilBERT), and fine-tuned sentence transformers. Forty-five classifier-embedding combinations were rigorously analyzed using accuracy, F1-score, and AUC-ROC metrics. The results reveal that base pre-trained transformer embeddings produce improved detection performance, with RoBERTa-base paired with a multi-layer perceptron classifier obtaining 98.93% accuracy and 98.87% F1-score. Contrary to predictions, fine-tuned sentence transformers underperformed base models, demonstrating that semantic similarity optimisation decreases phishing-relevant information. Traditional approaches produced competitive performance (98.11% accuracy) with much decreased processing requirements. These findings give empirical insight for embedding selection in phishing detection systems, revealing that base transformer models offer best performance without fine-tuning overhead, while classic techniques remain viable for resource-constrained deployments.
Actions (login required)
![]() |
View Item |
Tools
Tools