Srinivasan, Ammaiappan (2025) A Comparative Study of Machine Learning and Lexicon-Based Methods for Sentiment Classification. Masters thesis, Dublin, National College of Ireland.
Preview |
PDF (Master of Science)
Download (719kB) | Preview |
Preview |
PDF (Configuration Manual)
Download (1MB) | Preview |
Abstract
This study is a comparison between machine learning and lexicon-based sentiment classification on Twitter data using Sentiment140 dataset of 1.6 million tweets. The paper has constructed an entire sentiment analysis pipeline including preprocessing, feature extraction, training of the model and evaluation. TF-IDF vectorization was done to clean, tokenize and lemmatize the text data which was then classified using four supervised models which included Naive Bayes, Logistic Regression, Support Vector Machine (SVM) and random forest. Accuracy, precision, recall and F1-score were used to measure the performance. Findings indicated that both Logistic Regression and Linear SVM scored the highest accuracy and F1-scores of about 78 and better than lexicon-based methods on contextual adaptability and predictive reliability. Even though lexicon-based methods were interpretable and cheap to compute, they could not handle informal and ambiguous language. The paper finds that linear machine learning fits offer intermediate between scalability and accuracy and that the next research should adopt hybrid or transformer-based systems to capture the contextual sentiment in greater detail.
Actions (login required)
![]() |
View Item |
Tools
Tools