Birje, Sonali Anil (2025) Comparative Evaluation of RNN Architectures and Embedding Strategies for Multi-Label Disease Classification from Clinical Narratives. Masters thesis, Dublin, National College of Ireland.
Preview |
PDF (Master of Science)
Download (956kB) | Preview |
Preview |
PDF (Configuration Manual)
Download (1MB) | Preview |
Abstract
Unstructured clinical discharge summaries represent a valuable source of clinically relevant information, although they produce challenges to automated processing due to large amounts of narrative information, inclusion of domain-specific terminologies, and high rates of multiple comorbid conditions. In this respect, conventional machine-learning methods like Support Vector Machines and logistic regression are limited by the fact that they rely on manually designed features, by their inability to model long-range interactions, and by overfitting high-dimensional, sparse data. This paper suggests a deep-learning model that uses three recurrent neural network structures, including Vanilla RNN, Gated Recurrent Unit (GRU), and Bidirectional Long Short-Term Memory (BiLSTM), to classify diseases in multi-label classification of discharge summaries. A synthetic dataset was created in a balanced way with 2,000 records, matching the International Classification of Diseases, Ninth Revision (ICD-9) diagnostic codes with clinical narrative templates thus ensuring the equal distribution over the fifteen illness categories and not compromising the confidentiality of a patient. The text was preprocessed and tokenized, noise removed, and sequences padded, and then represented using two types of embeddings, general-purpose GloVe embeddings and domain-specific BioBERT contextual embeddings. All model-embedding combinations were trained with binary cross-entropy loss and early stopping rules and performance assessed on a reserved test set in terms of Micro and Macro F1-scores, precision, recall, and ROC-AUC. The results indicate that the BiLSTM model augmented with BioBERT embeddings has reached the best overall performance, as it produced a mean ROC-AUC of 0.92 and better F1-scores of most labels, therefore, supporting the idea of combining the context modeling capacity with domain-specific embeddings. These findings offer a reproducible framework of building ethical, accurate and interpretable natural-language processing applications in healthcare without relying on actual patient data.
| Item Type: | Thesis (Masters) |
|---|---|
| Supervisors: | Name Email Rustam, Furqan UNSPECIFIED |
| Uncontrolled Keywords: | Multi-label Classification; Discharge Summaries; RNN; GRU; BiLSTM; BioBERT; GloVe; Clinical NLP |
| Subjects: | P Language and Literature > P Philology. Linguistics > Computational linguistics. Natural language processing R Medicine > Healthcare Industry Q Science > Q Science (General) > Self-organizing systems. Conscious automata > Machine learning |
| Divisions: | School of Computing > Master of Science in Data Analytics |
| Depositing User: | Ciara O'Brien |
| Date Deposited: | 25 Aug 2026 11:39 |
| Last Modified: | 25 Aug 2026 11:39 |
| URI: | https://norma.ncirl.ie/id/eprint/9624 |
Actions (login required)
![]() |
View Item |
Tools
Tools