NORMA eResearch @NCI Library

Hybrid SHAP-LRP for Adversarial Debugging in Vision Transformers: A Layer-Wise Explainability Approach

Sawji, Shreyas Mukund (2025) Hybrid SHAP-LRP for Adversarial Debugging in Vision Transformers: A Layer-Wise Explainability Approach. Masters thesis, Dublin, National College of Ireland.

[thumbnail of Master of Science]
Preview
PDF (Master of Science)
Download (1MB) | Preview
[thumbnail of Configuration Manual]
Preview
PDF (Configuration Manual)
Download (717kB) | Preview

Abstract

Vision Transformers (ViTs) have been shown to perform well on image classification tasks but tend to be difficult to interpret and susceptible to adversarial attacks where small, often imperceptible, changes to input images result in incorrect predictions. This paper addresses the urgent needs of interpretability and robustness in ViTs through a new hybrid model that integrates SHAP (Shapley Additive Explanations) and Layer-wise Relevance Propagation (LRP) to explain and to diagnose adversarial vulnerability.

We trained a ViT-tiny Patch16-224 model on the CIFAR-10 dataset and applied adversarial attacks created by the Fast Gradient Sign Method (FGSM) to it. Our solution paired SHAP to map global feature importance and LRP to determine which internal transformer blocks were most impacted by attacks. Results revealed that adversarial perturbations caused obvious changes in SHAP attributions, and the model's attention shifted away from semantically important areas. At the same time, LRP manifested irregular behaviour in early transformer layers, indicating internal instability. By tracing those patterns, our method effectively identified not only misclassifications but also internal, subtle perturbations caused by adversarial inputs transmitting real-time alerts upon assessment.

This paper demonstrates that combining input and layer-level attribution methods provides actionable interpretability and improves the safety profile of ViT models. Even though the current work is evaluated on CIFAR-10 and FGSM attacks, future work will apply this hybrid mechanism to more advanced datasets, more powerful adversarial environments, and real-world applications such as diagnosis of medical imaging and autonomous systems.

Item Type: Thesis (Masters)
Supervisors:
Name
Email
Abgaz, Yalemisew
UNSPECIFIED
Uncontrolled Keywords: Neural Style Transfer; Generative Adversarial Networks; CycleGAN; Artistic Style Transfer; VGG19; Image-to-Image Translation; FID; SSIM; Deep Learning; PyTorchIntroduction (The Growing Energy Demand of Big Data Analytics)
Subjects: Q Science > QA Mathematics > Electronic computers. Computer science
T Technology > T Technology (General) > Information Technology > Electronic computers. Computer science
Q Science > QH Natural history > QH301 Biology > Methods of research. Technique. Experimental biology > Data processing. Bioinformatics > Artificial intelligence > Computer vision
Q Science > Q Science (General) > Self-organizing systems. Conscious automata > Artificial intelligence > Computer vision
Divisions: School of Computing > Master of Science in Data Analytics
Depositing User: Ciara O'Brien
Date Deposited: 26 Aug 2026 11:46
Last Modified: 26 Aug 2026 11:46
URI: https://norma.ncirl.ie/id/eprint/9666

Actions (login required)

View Item View Item