Gurrala, Ajay (2025) Fine tuning GPT- 3.5 for hate speech detection with continuous severity scoring. Masters thesis, Dublin, National College of Ireland.
Preview |
PDF (Master of Science)
Download (1MB) | Preview |
Preview |
PDF (Configuration Manual)
Download (1MB) | Preview |
Abstract
The spread of the harmful content in the social media platforms requires automated detection systems, which can recognize the severities in a nuanced manner that goes beyond binary classification. The study presents a very significant gap because it is the first supervised GPT-3.5 fine-tuning on continuous hate speech severity data, based on the UC Berkeley Measuring Hate Speech corpus. An innovative three-way classification scheme (HATE/ ABSTAIN/ NOTHATE) with score-enriched output mode was adopted, which allowed categorical prediction, as well as continuous estimation of severity at the same time. The fine-tuned model was compared with three base models BiLSTM, BERT-base and RoBERTa-base--with the same data partitions of 26, 557 training samples and 11, 382 test samples. The experimental findings show that fine-tuned GPT-3.5 is more efficient in all metrics: 70.68% classification accuracy (compared with 68.25 in RoBERTa-base), 0.6882 macro F1-score, 0.8862 mean absolute error and 0.8390 Pearson correlation of prediction of severity. Analysis of errors showed that the critical cases of HATE-NOTHATE misclassification were only found in 1.94% of the samples. The study can add a new dual-task learning model that keeps the continuous severity information that binary-based methods usually ignore, and it has an impact on content moderation systems that need to make both categorical and severity-based prioritization decisions.
Actions (login required)
![]() |
View Item |
Tools
Tools