NORMA eResearch @NCI Library

Reasoning-Enhanced Contrastive Learning for Detecting Hate Speech in LLM-Generated Content: A ModernBERT-Based Approach with Chain-of-Thought Supervision

Krishnakumari, Ashwin Jayakumar (2025) Reasoning-Enhanced Contrastive Learning for Detecting Hate Speech in LLM-Generated Content: A ModernBERT-Based Approach with Chain-of-Thought Supervision. Masters thesis, Dublin, National College of Ireland.

[thumbnail of Master of Science]
Preview
PDF (Master of Science)
Download (857kB) | Preview
[thumbnail of Configuration Manual]
Preview
PDF (Configuration Manual)
Download (762kB) | Preview

Abstract

When used on content made by Large Language Models, current hate speech detectors show a big drop in performance, with accuracy dropping by 10 to 15 percentage points compared to text written by people. This degradation is due to relying on keywords, not being able to see hidden patterns, not being able to see the whole picture, and being sensitive to differences in style across LLM sources. This study examines the potential of reasoning-enhanced contrastive learning to bridge the detection gap by integrating explicit chain-of-thought reasoning with long-context encoding. The suggested framework combines three new parts: OpenAI's o3-mini model for making detailed analytical reasoning chains, ModernBERT's 8,192-token context window for processing full reasoning without cutting it off, and a multi-objective training method that uses classification, contrastive, and alignment losses. The HateBenchSet dataset, which has 7,838 samples from six LLMs, shows that the proposed approach gets an F1 score of 98.87%, which is 7.46 points better than the baseline and 12.87 points better than the best existing detector. Ablation studies show that reasoning chains are the most important factor, making things 8.50% better, while contrastive learning has a small effect. The study introduces an innovative detection framework, marks the inaugural application of o3-mini and ModernBERT in hate speech detection, and provides a new dataset resource featuring reasoning-annotated samples for subsequent research on explainable content moderation.

Item Type: Thesis (Masters)
Supervisors:
Name
Email
Tomer, Vikas
UNSPECIFIED
Subjects: Q Science > QH Natural history > QH301 Biology > Methods of research. Technique. Experimental biology > Data processing. Bioinformatics > Artificial intelligence
Q Science > Q Science (General) > Self-organizing systems. Conscious automata > Artificial intelligence
Z Bibliography. Library Science. Information Resources > ZA Information resources > ZA4150 Computer Network Resources > The Internet > World Wide Web > Websites > Online social networks
T Technology > TK Electrical engineering. Electronics. Nuclear engineering > Telecommunications > The Internet > World Wide Web > Websites > Online social networks
Divisions: School of Computing > Master of Science in Data Analytics
Depositing User: Ciara O'Brien
Date Deposited: 08 Sep 2026 08:47
Last Modified: 08 Sep 2026 08:47
URI: https://norma.ncirl.ie/id/eprint/9873

Actions (login required)

View Item View Item