NORMA eResearch @NCI Library

Hybrid Audio Event Recognition Using CNN and Transfer Learning: Comparative Study on Custom and Google Speech Commands Datasets

Kumthekar, Hrushikesh Nitin (2025) Hybrid Audio Event Recognition Using CNN and Transfer Learning: Comparative Study on Custom and Google Speech Commands Datasets. Masters thesis, Dublin, National College of Ireland.

[thumbnail of Master of Science]
Preview
PDF (Master of Science)
Download (1MB) | Preview
[thumbnail of Configuration Manual]
Preview
PDF (Configuration Manual)
Download (509kB) | Preview

Abstract

The current project deeply analyses how data size emphasizes on reliability as well as reliability in field of keyword spotting (KWS) models for detection of distress words related audio circumstances. Phrases like “save me” or “help me” as well as gunshot shots were identified moreover is also vital for crucial for execution of real-time emergency response systems. Typically, deep learning models are usually effective when trained on smaller size datasets and how model performance is influenced by data scale and quality. Two datasets were used for this research. A custom-made dataset of approximately of size ~300 .wav audio samples were generated using text-to-speech tools and public sound libraries to simulate emergency situations. Concurrently, the large dataset Google Speech Commands (GSC) comprising of roughly 65,000 audio samples was used to benchmark the executed system. The implemented two CNN and pre-trained YAMNet models with classifiers like SVM were trained and evaluated across these two used datasets. Mel spectrograms being standard pre-processing technique was being applied in both cases thus revealing results which showed clear trend. The custom-made CNN achieved high to moderate accuracy across both datasets but ineffective to generalize correctly. Although model demonstrated moderate accuracy predictions were inaccurate frequently. On the other hand, GSC custom trained CNN model achieved higher accuracy and was able to produce more correct predictions which was reliable during real-time testing. This finding emphasizes the limitations of solely reliant upon evaluation metrics and highlights usage of deeper value of large, vast dataset in build-up of robust models. This end results suggest that quality and size of data not only impact accuracy but also actual prediction capabilities. Future work focuses on improvement in variability of small datasets and realism though augmentation techniques, collection of real-world audio samples and validating models within more complex and noisy environments to ensure dependable KWS system.

Item Type: Thesis (Masters)
Supervisors:
Name
Email
Nolan, Eamon
UNSPECIFIED
Uncontrolled Keywords: Keyword Spotting; Dataset Size; Distress Detection; Emergency Audio; Deep Learning; Small Dataset; Model Reliability; CNN; YAMNet; SVM; Mel Spectrogram; Data Augmentation; Real-world Testing; Dataset Quality
Subjects: Q Science > QA Mathematics > Electronic computers. Computer science
T Technology > T Technology (General) > Information Technology > Electronic computers. Computer science
H Social Sciences > HM Sociology > Information Science > Communication
Divisions: School of Computing > Master of Science in Data Analytics
Depositing User: Ciara O'Brien
Date Deposited: 25 Aug 2026 15:51
Last Modified: 25 Aug 2026 15:51
URI: https://norma.ncirl.ie/id/eprint/9641

Actions (login required)

View Item View Item