Jiang, Boyang (2025) A Machine Learning-based Monitoring Framework for Prediction, Diagnosis, and Feedback Learning in DevOps. Masters thesis, Dublin, National College of Ireland.
Preview |
PDF (Master of Science)
Download (1MB) | Preview |
Preview |
PDF (Configuration Manual)
Download (981kB) | Preview |
Abstract
Modern application complexity often leads to faults detected only after user impact. This paper reflects on real-world experiences to identify persistent challenges such as fault prediction, root cause diagnosis, and post-failure learning in enterprise monitoring systems. To address these issues, we designed machine learning monitoring framework through a closed loop based on the Google SRE’s incident lifecycle model: pre-failure, in-failure, and post-failure to enhance detection accuracy, reduce MTTR (i.e., mean time to recovery), and improve system reliability. Through a review of existing literature, we confirm that these issues are widely acknowledged across academia and industry. We further summarize the common limitations of current solutions and outline key technical challenges. Finally, this proposal outlines a hypothetical methodology that uses various data sources (e.g., Log, DevOps Pipeline) to train the model. These models are mapped to each lifecycle stage, followed by a Gantt chart outlining the implementation plan.
| Item Type: | Thesis (Masters) |
|---|---|
| Supervisors: | Name Email Heeney, Sean UNSPECIFIED |
| Subjects: | T Technology > T Technology (General) > Information Technology > Cloud computing Q Science > Q Science (General) > Self-organizing systems. Conscious automata > Machine learning |
| Divisions: | School of Computing > Master of Science in Cloud Computing |
| Depositing User: | Ciara O'Brien |
| Date Deposited: | 31 Aug 2026 14:49 |
| Last Modified: | 31 Aug 2026 14:49 |
| URI: | https://norma.ncirl.ie/id/eprint/9707 |
Actions (login required)
![]() |
View Item |
Tools
Tools