NORMA eResearch @NCI Library

NLP and Machine Learning Framework for Reducing False Positive in Watchlist Name Screening in AML Systems

Jadhav, Gayatri Mahesh (2025) NLP and Machine Learning Framework for Reducing False Positive in Watchlist Name Screening in AML Systems. Masters thesis, Dublin, National College of Ireland.

[thumbnail of Master of Science]
Preview
PDF (Master of Science)
Download (946kB) | Preview
[thumbnail of Configuration Manual]
Preview
PDF (Configuration Manual)
Download (654kB) | Preview

Abstract

The Anti-Money laundering (AML) systems screening sanctions often give high false-positive rates as a result of linguistic variation, lack of consistency in transliteration and typing errors. Conventional rule-based and fuzzy string-matching algorithms, e.g. Levenshtein and Jaro-Winkler, typically have a hard time representing semantic relationships between names, leading to higher operational costs and regulatory inefficiencies.

This study created a hybrid Natural Language Processing (NLP) and Machine Learning (ML) model that combines Sentence-BERT (SBERT) semantic embeddings with traditional string-similarity scores with the help of Shapley Additive Explanations (SHAP) to make the results easier to interpret. Synthetic customer records were generated using standard data-generation techniques to maintain privacy. Random Forest, XGBoost and Support Vector Machine (SVM) which are three supervised classifiers were trained on the integrated semantic-fuzzy feature set and assessed in terms of Precision, Recall, F1-score, ROC-AUC and False Positive Rate.

The discrimination of matching and non-matching name pairs was moderate in all models. SVM had the best generalisation with the highest ROC-AUC score, and the other two models, Random Forest and XGBoost, had more recall at the expense of more false positives. SHAP analysis revealed that fuzzy similarity measures, especially TokenSetRatio, had the strongest impact on all models, and the contribution of features was in line with human intuition.

In theory, the findings demonstrate that contextual embeddings can be used to supplement surface-level similarity in short-text name-matching tasks. In practice, the hybrid framework will increase the transparency, interpretability and auditability of AML screening processes, which is in line with regulatory requirements of explainable AI. More advancements in predictive accuracy will need more multilingual and phonetic and relational features and access to real operational data.

Item Type: Thesis (Masters)
Supervisors:
Name
Email
Niculescu, Hamilton
UNSPECIFIED
Subjects: H Social Sciences > HG Finance
P Language and Literature > P Philology. Linguistics > Computational linguistics. Natural language processing
Q Science > Q Science (General) > Self-organizing systems. Conscious automata > Machine learning
Divisions: School of Computing > Master of Science in Data Analytics
Depositing User: Ciara O'Brien
Date Deposited: 07 Sep 2026 11:52
Last Modified: 07 Sep 2026 11:52
URI: https://norma.ncirl.ie/id/eprint/9864

Actions (login required)

View Item View Item