NORMA eResearch @NCI Library

Vision-Language Model Fine-Tuning for Face Mask Detection Using Parameter-Efficient Methods

Gollapudi, Chandrika (2025) Vision-Language Model Fine-Tuning for Face Mask Detection Using Parameter-Efficient Methods. Masters thesis, Dublin, National College of Ireland.

[thumbnail of Master of Science]
Preview
PDF (Master of Science)
Download (1MB) | Preview
[thumbnail of Configuration Manual]
Preview
PDF (Configuration Manual)
Download (764kB) | Preview

Abstract

During the COVID-19 pandemic, it became very important for public health to keep an eye on people who wear face masks. However, the current detection systems use traditional convolutional neural networks that only give bounding boxes and class labels without any clear explanations. This study examines the potential of vision-language models, fine-tuned using parameter-efficient techniques, to attain competitive detection accuracy while providing improved interpretability through natural language outputs. The research assesses two parameter-efficient fine-tuning methodologies, Low-Rank Adaptation and Quantized Low-Rank Adaptation, implemented on Microsoft's Florence-2 vision-language model for face mask detection across three compliance categories. Experiments employed an extensive dataset comprising 28,297 images and 63,035 annotations, characterized by a pronounced class imbalance. The LoRA-adapted Florence-2-large reached 71.6% mAP@50 with 88.4% precision and 87.3% recall while training only 0.7% of model parameters. This was higher than the seventy percent accuracy level that was set as the primary success level. QLoRA setting achieved 40.1 percent mAP at 50 indicating that aggressive quantizing four-bit quantization has a significant impact on detecting vision tasks that require accurate spatial positioning. Vision-language models demonstrate that they can attain realistic detection accuracy to compliance monitoring uses as well as equip output understandable building blocks when compared to a YOLO11m baseline (95.6% mAP@50). The findings indicate that finetuning using a small number of parameters is an appropriate method to harmone large vision-language models to safety monitoring duties in a particular area and does not require a significant amount of computing hardware.

Item Type: Thesis (Masters)
Supervisors:
Name
Email
Garg, Mohit
UNSPECIFIED
Subjects: R Medicine > RA Public aspects of medicine > RA0421 Public health. Hygiene. Preventive Medicine
Q Science > QH Natural history > QH301 Biology > Methods of research. Technique. Experimental biology > Data processing. Bioinformatics > Artificial intelligence
Q Science > Q Science (General) > Self-organizing systems. Conscious automata > Artificial intelligence
P Language and Literature > P Philology. Linguistics > Computational linguistics. Natural language processing
Q Science > QH Natural history > QH301 Biology > Methods of research. Technique. Experimental biology > Data processing. Bioinformatics > Artificial intelligence > Computer vision
Q Science > Q Science (General) > Self-organizing systems. Conscious automata > Artificial intelligence > Computer vision
R Medicine > Diseases > Outbreaks of disease > Epidemics > COVID-19 Pandemic, 2020-
Divisions: School of Computing > Master of Science in Artificial Intelligence
Depositing User: Ciara O'Brien
Date Deposited: 02 Sep 2026 09:16
Last Modified: 02 Sep 2026 09:16
URI: https://norma.ncirl.ie/id/eprint/9755

Actions (login required)

View Item View Item