NORMA eResearch @NCI Library

Adaptive Safety Moderation for Large Language Models: A Context-Aware Approach to Mitigating Adversarial Prompts

Iftikhar, Adnan (2025) Adaptive Safety Moderation for Large Language Models: A Context-Aware Approach to Mitigating Adversarial Prompts. Masters thesis, Dublin, National College of Ireland.

[thumbnail of Master of Science]
Preview
PDF (Master of Science)
Download (1MB) | Preview
[thumbnail of Configuration Manual]
Preview
PDF (Configuration Manual)
Download (1MB) | Preview

Abstract

Safety for large language models is ultimately a system problem, not a single-classifier problem. We evaluate a practical moderation pipeline that composes lightweight routing, specialised experts (e.g., hate, self-harm, misinformation, PII), calibrated aggregation into a global risk score, and an explicit policy that maps scores to actions (allow / clarify / redact / block). A calibrated LinearSVC (TF-IDF) serves as our baseline. On a held-out test set (N=5,175), the baseline attains accuracy 0.795, macro-F1 0.795, and AP (unsafe) ≈ 0.863, yet at τ=0.50 still yields FN 587 and FP 472 (jailbreak ≈ 22.3%, over-refusal ≈ 18.5%). Qualitative checks echo the aggregates: a benign query is allowed; explicit violent intent is blocked; but an overt hate statement near the decision boundary is incorrectly passed as safe. We therefore frame safety optimisation as policy tuning on calibrated scores and augment the system with (i) category-aware routing, (ii) a severity-aware override in the mid-band to prefer clarify/redact over allow, and (iii) auditable rewrite tags for PII. Evaluation focuses on detector quality and policy outcomes (jailbreak vs over-refusal), with ablations for routing, calibration, and thresholds. The approach offers deployers actionable levers—thresholds tied to risk appetite, per-category signals, and transparent rationales—while acknowledging limits (domain shift, multilingual coverage). Future work targets threat-model-aware benchmarking at scale and category-specific calibration to further reduce harmful misses without unacceptable refusals.

Item Type: Thesis (Masters)
Supervisors:
Name
Email
Makki, Ahmed
UNSPECIFIED
Subjects: Q Science > QA Mathematics > Electronic computers. Computer science
T Technology > T Technology (General) > Information Technology > Electronic computers. Computer science
Q Science > QH Natural history > QH301 Biology > Methods of research. Technique. Experimental biology > Data processing. Bioinformatics > Artificial intelligence
Q Science > Q Science (General) > Self-organizing systems. Conscious automata > Artificial intelligence
Divisions: School of Computing > Master of Science in Data Analytics
Depositing User: Ciara O'Brien
Date Deposited: 25 Aug 2026 14:46
Last Modified: 25 Aug 2026 14:46
URI: https://norma.ncirl.ie/id/eprint/9631

Actions (login required)

View Item View Item