Jadhav, Gayatri Mahesh (2025) NLP and Machine Learning Framework for Reducing False Positive in Watchlist Name Screening in AML Systems. Masters thesis, Dublin, National College of Ireland.
Preview |
PDF (Master of Science)
Download (946kB) | Preview |
Preview |
PDF (Configuration Manual)
Download (654kB) | Preview |
Abstract
The Anti-Money laundering (AML) systems screening sanctions often give high false-positive rates as a result of linguistic variation, lack of consistency in transliteration and typing errors. Conventional rule-based and fuzzy string-matching algorithms, e.g. Levenshtein and Jaro-Winkler, typically have a hard time representing semantic relationships between names, leading to higher operational costs and regulatory inefficiencies.
This study created a hybrid Natural Language Processing (NLP) and Machine Learning (ML) model that combines Sentence-BERT (SBERT) semantic embeddings with traditional string-similarity scores with the help of Shapley Additive Explanations (SHAP) to make the results easier to interpret. Synthetic customer records were generated using standard data-generation techniques to maintain privacy. Random Forest, XGBoost and Support Vector Machine (SVM) which are three supervised classifiers were trained on the integrated semantic-fuzzy feature set and assessed in terms of Precision, Recall, F1-score, ROC-AUC and False Positive Rate.
The discrimination of matching and non-matching name pairs was moderate in all models. SVM had the best generalisation with the highest ROC-AUC score, and the other two models, Random Forest and XGBoost, had more recall at the expense of more false positives. SHAP analysis revealed that fuzzy similarity measures, especially TokenSetRatio, had the strongest impact on all models, and the contribution of features was in line with human intuition.
In theory, the findings demonstrate that contextual embeddings can be used to supplement surface-level similarity in short-text name-matching tasks. In practice, the hybrid framework will increase the transparency, interpretability and auditability of AML screening processes, which is in line with regulatory requirements of explainable AI. More advancements in predictive accuracy will need more multilingual and phonetic and relational features and access to real operational data.
| Item Type: | Thesis (Masters) |
|---|---|
| Supervisors: | Name Email Niculescu, Hamilton UNSPECIFIED |
| Subjects: | H Social Sciences > HG Finance P Language and Literature > P Philology. Linguistics > Computational linguistics. Natural language processing Q Science > Q Science (General) > Self-organizing systems. Conscious automata > Machine learning |
| Divisions: | School of Computing > Master of Science in Data Analytics |
| Depositing User: | Ciara O'Brien |
| Date Deposited: | 07 Sep 2026 11:52 |
| Last Modified: | 07 Sep 2026 11:52 |
| URI: | https://norma.ncirl.ie/id/eprint/9864 |
Actions (login required)
![]() |
View Item |
Tools
Tools