NORMA eResearch @NCI Library

Embedding-Based Deep Learning Framework for Detecting Prompt Injection Attacks in Large Language Models

Palampalle, Padmavati (2025) Embedding-Based Deep Learning Framework for Detecting Prompt Injection Attacks in Large Language Models. Masters thesis, Dublin, National College of Ireland.

[thumbnail of Master of Science]
Preview
PDF (Master of Science)
Download (953kB) | Preview
[thumbnail of Configuration Manual]
Preview
PDF (Configuration Manual)
Download (869kB) | Preview

Abstract

Prompt injection attacks have serious security threats to Large Language Models (LLMs) by embedding malicious instructions within benign-looking prompts, causing models to execute unauthorized actions or leak sensitive information. There are some traditional rule-based and keyword detection mechanisms which fail against like complex, context-aware manipulations, intelligent adaptive detection systems. This study addresses the important need for strong prompt injection detection by developing and evaluating embedding-based deep learning classifiers capable of understanding semantic intent rather than trusting on surface-level patterns. Using the Malicious Prompts dataset from Hugging Face containing 373,646 records this study have implemented and compared five architectures which includes LSTM and BiLSTM models with FastText and GloVe embeddings, using a novel BiLSTM with Scalar Self-Attention Layer (SSAL). The proposed BiLSTM+SSAL architecture achieved superior performance with 0.82 accuracy, precision, recall and F1-score, outperforming traditional LSTM and standard BiLSTM models. The attention mechanism enabled dynamic weighting of contextually relevant tokens reducing false negatives from 2426 to 1718 compared to standard BiLSTM with GloVe. A Flask-based web interface combined with Groq LLM shows real-time classification capability. This study advances LLM security by establishing embedding-based classifiers though challenges remain in multilingual generalization and detecting growing adversarial tactics. Future work will explore transformer-based architectures and continual learning for adaptive threat detection.

Item Type: Thesis (Masters)
Supervisors:
Name
Email
Mahajan, Kamil
UNSPECIFIED
Uncontrolled Keywords: Prompt Injection Attacks; Large Language Models; Embedding-Based Classifiers; BiLSTM with Self-Attention
Subjects: Q Science > QH Natural history > QH301 Biology > Methods of research. Technique. Experimental biology > Data processing. Bioinformatics > Artificial intelligence
Q Science > Q Science (General) > Self-organizing systems. Conscious automata > Artificial intelligence
Q Science > QA Mathematics > Computer software > Computer Security
T Technology > T Technology (General) > Information Technology > Computer software > Computer Security
Divisions: School of Computing > Master of Science in Cyber Security
Depositing User: Ciara O'Brien
Date Deposited: 04 Sep 2026 09:07
Last Modified: 04 Sep 2026 09:07
URI: https://norma.ncirl.ie/id/eprint/9816

Actions (login required)

View Item View Item