Iftikhar, Adnan (2025) Adaptive Safety Moderation for Large Language Models: A Context-Aware Approach to Mitigating Adversarial Prompts. Masters thesis, Dublin, National College of Ireland.
Preview |
PDF (Master of Science)
Download (1MB) | Preview |
Preview |
PDF (Configuration Manual)
Download (1MB) | Preview |
Abstract
Safety for large language models is ultimately a system problem, not a single-classifier problem. We evaluate a practical moderation pipeline that composes lightweight routing, specialised experts (e.g., hate, self-harm, misinformation, PII), calibrated aggregation into a global risk score, and an explicit policy that maps scores to actions (allow / clarify / redact / block). A calibrated LinearSVC (TF-IDF) serves as our baseline. On a held-out test set (N=5,175), the baseline attains accuracy 0.795, macro-F1 0.795, and AP (unsafe) ≈ 0.863, yet at τ=0.50 still yields FN 587 and FP 472 (jailbreak ≈ 22.3%, over-refusal ≈ 18.5%). Qualitative checks echo the aggregates: a benign query is allowed; explicit violent intent is blocked; but an overt hate statement near the decision boundary is incorrectly passed as safe. We therefore frame safety optimisation as policy tuning on calibrated scores and augment the system with (i) category-aware routing, (ii) a severity-aware override in the mid-band to prefer clarify/redact over allow, and (iii) auditable rewrite tags for PII. Evaluation focuses on detector quality and policy outcomes (jailbreak vs over-refusal), with ablations for routing, calibration, and thresholds. The approach offers deployers actionable levers—thresholds tied to risk appetite, per-category signals, and transparent rationales—while acknowledging limits (domain shift, multilingual coverage). Future work targets threat-model-aware benchmarking at scale and category-specific calibration to further reduce harmful misses without unacceptable refusals.
| Item Type: | Thesis (Masters) |
|---|---|
| Supervisors: | Name Email Makki, Ahmed UNSPECIFIED |
| Subjects: | Q Science > QA Mathematics > Electronic computers. Computer science T Technology > T Technology (General) > Information Technology > Electronic computers. Computer science Q Science > QH Natural history > QH301 Biology > Methods of research. Technique. Experimental biology > Data processing. Bioinformatics > Artificial intelligence Q Science > Q Science (General) > Self-organizing systems. Conscious automata > Artificial intelligence |
| Divisions: | School of Computing > Master of Science in Data Analytics |
| Depositing User: | Ciara O'Brien |
| Date Deposited: | 25 Aug 2026 14:46 |
| Last Modified: | 25 Aug 2026 14:46 |
| URI: | https://norma.ncirl.ie/id/eprint/9631 |
Actions (login required)
![]() |
View Item |
Tools
Tools