White, Emily Elizabeth (2025) Evaluating the Ethical Quality of Generative AI Responses in Mental Health Contexts using a Constitutional AI Framework. Masters thesis, Dublin, National College of Ireland.
Preview |
PDF (Master of Science)
Download (1MB) | Preview |
Preview |
PDF (Configuration Manual)
Download (671kB) | Preview |
Abstract
The increasing use of large language models (LLMs) for mental health related queries raises important questions about ethical reliability, user safety and the consistency of model behaviour across providers. While European and UK professional bodies have established comprehensive ethical frameworks, practical tools for evaluating and improving LLM ethical performance remain limited. This study addresses that gap by developing a Weighted Ethical Evaluation Rubric grounded in European and UK psychological and AI Act Safety principles and by testing whether a Constitutional AI (CAI) style prompt intervention can improve the ethical quality of LLM responses.
A synthetic dataset of 1020 mental health prompts was generated and evaluated across three leading LLMs: GPT-4o (OpenAI), Claude Sonnet 4 (Anthropic) and Gemini 2.5 Pro (Google), under two conditions: a Baseline condition with no ethical guidance and a CAI Prefix condition in which prompts were preceded by an ethical instruction. This produced 6120 model responses, each scored using the rubric and categorised into ethical tiers. Statistical analysis, including Wilcoxon signed-rank tests and effect-size measures, compared ethical quality across models and conditions.
Results show that the CAI Prefix significantly improved ethical scores for all models, with GPT-4o demonstrating the largest improvement (Δ = 0.62, very large effect). Claude and Gemini showed smaller but statistically strong gains. These findings indicate that prompt level ethical instruction can shape how LLMs respond to sensitive mental health prompts, particularly for models without embedded constitutional training.
This research contributes a practical, reproducible methodology for assessing LLM ethical behaviour and demonstrates how professional ethical frameworks can be operationalised into measurable evaluation criteria. The study highlights the potential for prompt based ethical guidance to enhance model safety in high risk domains and provides a foundation for future work on conversational alignment, expanded rubric dimensions and real world deployment of ethical guardrails.
| Item Type: | Thesis (Masters) |
|---|---|
| Supervisors: | Name Email Del Rosal, Victor UNSPECIFIED |
| Uncontrolled Keywords: | Constitutional AI; AI ethics; mental health; CAI Prefix; ethical quality |
| Subjects: | B Philosophy. Psychology. Religion > BJ Ethics Q Science > QH Natural history > QH301 Biology > Methods of research. Technique. Experimental biology > Data processing. Bioinformatics > Artificial intelligence > Generative artificial intelligence Q Science > Q Science (General) > Self-organizing systems. Conscious automata > Artificial intelligence > Generative artificial intelligence R Medicine > RA Public aspects of medicine > RA790 Mental Health |
| Divisions: | School of Computing > Master of Science in Artificial Intelligence for Business |
| Depositing User: | Ciara O'Brien |
| Date Deposited: | 03 Sep 2026 08:48 |
| Last Modified: | 03 Sep 2026 08:48 |
| URI: | https://norma.ncirl.ie/id/eprint/9783 |
Actions (login required)
![]() |
View Item |
Tools
Tools