Singh, Utkarsh (2025) Performance Optimization and Cost Analysis of qpAdm Genomic Modeling on AWS Cloud Infrastructure. Masters thesis, Dublin, National College of Ireland.
Preview |
PDF (Master of Science)
Download (1MB) | Preview |
Preview |
PDF (Configuration Manual)
Download (728kB) | Preview |
Abstract
In the rapidly evolving field of paleogenomics, the analysis of ancient DNA (aDNA) has become a cornerstone of understanding human population history. Genomic analyses increasingly rely on computationally intensive workflows that integrate large genotype matrices with sophisticated statistical methods. Among these, qpAdm and related f-statistics tools from the ADMIXTOOLS suite are central to modelling population admixture, yet there is limited guidance on how to deploy them efficiently and reproducibly in cloud environments. This thesis investigates how to optimise the cost–performance trade-offs of qpAdm workflows on Amazon Web Services (AWS), while preserving reproducibility and scalability. A fully containerised pipeline was developed that encapsulates ADMIXTOOLS, PLINK and associated utilities within a version-pinned Docker image, deployed on EC2 instances provisioned via Terraform and driven by a lightweight benchmark runner inspired by Snakemake. Three representative datasets were used: a subset of the Allen Ancient DNA Resource (AADR) 1240k panel, a South Asian subset of the 1000 Genomes Project on chromosome 22, and a smaller synthetic dataset derived from SGDP-like structure. Experiments evaluated Graviton-based c7g, m7g and r7g instance families under on-demand and Spot pricing, with multiple repetitions per configuration and statistical analysis of log-transformed runtimes. The results show that preprocessing and format conversion dominate end-to-end runtime for non-EIGENSTRAT data, while qpAdm and qpfstats computations are primarily CPU-bound and scale well on compute-optimised Graviton3 hardware. c7g instances consistently provided the best price–performance among the evaluated families, and Spot instances reduced estimated compute costs by around two-thirds without compromising analytical validity. For the full AADR v62 1240k panel, we observed qpfstats consuming more than 31 GiB of memory and triggering out-of-memory termination on r7g.xlarge, motivating the use of r7g.2xlarge for the largest workloads. The pipeline, together with the empirical findings and derived heuristics, provides a practical blueprint for cost-effective, reproducible qpAdm modelling on AWS.
| Item Type: | Thesis (Masters) |
|---|---|
| Supervisors: | Name Email Emani, Sai UNSPECIFIED |
| Uncontrolled Keywords: | ancient DNA; population admixture; qpAdm; f-statistics; ADMIXTOOLS; Amazon Web Services (AWS); cloud benchmarking; Graviton3; EC2; Spot Instances; Docker; Terraform; reproducible workflows; preprocessing; EIGENSTRAT; PLINK; cost–performance optimisation |
| Subjects: | Q Science > QA Mathematics > Electronic computers. Computer science T Technology > T Technology (General) > Information Technology > Electronic computers. Computer science T Technology > T Technology (General) > Information Technology > Cloud computing |
| Divisions: | School of Computing > Master of Science in Cloud Computing |
| Depositing User: | Ciara O'Brien |
| Date Deposited: | 01 Sep 2026 11:10 |
| Last Modified: | 01 Sep 2026 11:10 |
| URI: | https://norma.ncirl.ie/id/eprint/9738 |
Actions (login required)
![]() |
View Item |
Tools
Tools