Kumthekar, Hrushikesh Nitin (2025) Hybrid Audio Event Recognition Using CNN and Transfer Learning: Comparative Study on Custom and Google Speech Commands Datasets. Masters thesis, Dublin, National College of Ireland.
Preview |
PDF (Master of Science)
Download (1MB) | Preview |
Preview |
PDF (Configuration Manual)
Download (509kB) | Preview |
Abstract
The current project deeply analyses how data size emphasizes on reliability as well as reliability in field of keyword spotting (KWS) models for detection of distress words related audio circumstances. Phrases like “save me” or “help me” as well as gunshot shots were identified moreover is also vital for crucial for execution of real-time emergency response systems. Typically, deep learning models are usually effective when trained on smaller size datasets and how model performance is influenced by data scale and quality. Two datasets were used for this research. A custom-made dataset of approximately of size ~300 .wav audio samples were generated using text-to-speech tools and public sound libraries to simulate emergency situations. Concurrently, the large dataset Google Speech Commands (GSC) comprising of roughly 65,000 audio samples was used to benchmark the executed system. The implemented two CNN and pre-trained YAMNet models with classifiers like SVM were trained and evaluated across these two used datasets. Mel spectrograms being standard pre-processing technique was being applied in both cases thus revealing results which showed clear trend. The custom-made CNN achieved high to moderate accuracy across both datasets but ineffective to generalize correctly. Although model demonstrated moderate accuracy predictions were inaccurate frequently. On the other hand, GSC custom trained CNN model achieved higher accuracy and was able to produce more correct predictions which was reliable during real-time testing. This finding emphasizes the limitations of solely reliant upon evaluation metrics and highlights usage of deeper value of large, vast dataset in build-up of robust models. This end results suggest that quality and size of data not only impact accuracy but also actual prediction capabilities. Future work focuses on improvement in variability of small datasets and realism though augmentation techniques, collection of real-world audio samples and validating models within more complex and noisy environments to ensure dependable KWS system.
| Item Type: | Thesis (Masters) |
|---|---|
| Supervisors: | Name Email Nolan, Eamon UNSPECIFIED |
| Uncontrolled Keywords: | Keyword Spotting; Dataset Size; Distress Detection; Emergency Audio; Deep Learning; Small Dataset; Model Reliability; CNN; YAMNet; SVM; Mel Spectrogram; Data Augmentation; Real-world Testing; Dataset Quality |
| Subjects: | Q Science > QA Mathematics > Electronic computers. Computer science T Technology > T Technology (General) > Information Technology > Electronic computers. Computer science H Social Sciences > HM Sociology > Information Science > Communication |
| Divisions: | School of Computing > Master of Science in Data Analytics |
| Depositing User: | Ciara O'Brien |
| Date Deposited: | 25 Aug 2026 15:51 |
| Last Modified: | 25 Aug 2026 15:51 |
| URI: | https://norma.ncirl.ie/id/eprint/9641 |
Actions (login required)
![]() |
View Item |
Tools
Tools