NORMA eResearch @NCI Library

Early recognition of Sepsis based on explainable time-series deep learning and Machine Learning models on real ICU data

Paul, Aleena (2025) Early recognition of Sepsis based on explainable time-series deep learning and Machine Learning models on real ICU data. Masters thesis, Dublin, National College of Ireland.

[thumbnail of Master of Science]
Preview
PDF (Master of Science)
Download (996kB) | Preview
[thumbnail of Configuration Manual]
Preview
PDF (Configuration Manual)
Download (352kB) | Preview

Abstract

Sepsis is a life-threatening and quickly evolving condition, which needs to be identified at the earliest stage, however, standardized criteria used to officially assess the condition such as SOFA and SIRS do not reflect the constantly changing time-dependent physiology of ICU patients, thus leading to the creation of data-intensive predictive models that may lack interpretability or generalisability. The paper builds and compares explainable time-series models predicting early sepsis based on the PhysioNet/Sepsis Challenge 2019 dataset and trains classical machine learning models (Logistic Regression, Random Forest, Gradient Boosting, XGBoost, LightGBM), models based on deep learning (LSTM, GRU, BiLSTM, CNN-1D) on rigorous preprocessing to deal with missingness and irregular sampling conditions. The evaluation of the performance was done based on the values of AUROC and AUPRC, F1, sensitivity, specificity, and calibration values. The best model was LightGBM whose AUROC is 0.877 that is stronger than XGBoost by 0.003, Gradient Boosting by 0.089, Random Forest by 0.025 and Logistic Regression by 0.039 and its discriminative ability is better than all other ML baselines. Deep learning models had much worse results with GRU having the highest AUROC of 0.623, showing that classical assembly is significantly more effective on this ICU dataset. SHAP explainability found the most effective predictors to be ICULOS, HospAdmTime, HR, and O2Sat. Calibration analysis indicated that LightGBM made reasonable probability estimates (Brier Score = 0.0979), and operating point analysis indicated that LightGBM could only give 0.701 specificity at 0.85 sensitivity (threshold = 0.2779), but GRU has only 0.199 specificity at the same constraint. In general, it can be concluded that tree-based ensembles, especially LightGBM, can offer good and explainable early-warning capabilities, and future studies will focus on enhancing temporal modelling and prospective validation of models in real-time ICU application.

Item Type: Thesis (Masters)
Supervisors:
Name
Email
Zahoor, Sheresh
UNSPECIFIED
Subjects: R Medicine > Healthcare Industry
Q Science > Q Science (General) > Self-organizing systems. Conscious automata > Machine learning
Divisions: School of Computing > Master of Science in Data Analytics
Depositing User: Ciara O'Brien
Date Deposited: 08 Sep 2026 11:31
Last Modified: 08 Sep 2026 11:31
URI: https://norma.ncirl.ie/id/eprint/9893

Actions (login required)

View Item View Item