Prudnikovas, Aivaras (2025) LLM based prompt injection detection accuracy and its relationship to the context size. Masters thesis, Dublin, National College of Ireland.
Preview |
PDF (Master of Science)
Download (1MB) | Preview |
Preview |
PDF (Configuration Manual)
Download (1MB) | Preview |
Abstract
LLMs, widely used in various applications, are vulnerable to malicious prompt injections that can manipulate their outputs, potentially leading to security breaches. Prompt injection attacks exploit the model’s design to follow instructions allowing attackers to change the outputs by injecting malicious instructions. This research evaluates how the prompt size influences the accuracy of the prompt injection attacks. It also evaluates the accuracy of two detection techniques (naïve and reinforced) on GPT-4.1, Ministral-3b, Mistral-large, Llama-4-Maverick, Gemma3. A set of more than 1000 injection prompts are wrapped in varying size content to increase context (no increase, 1000 tokens, 5000 tokens, 10000 tokens, 50000 tokens) and detection performance is measured for each configuration. The results show the detection accuracy first increases when context size is inflated with 1000 tokens of HTML on all LLMs and detections, then it varies without a clear relationship. GPT-4.1 outperforms other models peaking at 98.8% accuracy and a reinforced detection outperforms the naïve one.
Actions (login required)
![]() |
View Item |
Tools
Tools