Swarna, Rajesh Chowdhary (2025) Integrated Multimodal AI Framework for Enhanced Misinformation Detection: Combining LSTM Text Classification, GPT-4o Vision Analysis, and Professional Fact-Checking. Masters thesis, Dublin, National College of Ireland.
Preview |
PDF (Master of Science)
Download (1MB) | Preview |
Preview |
PDF (Configuration Manual)
Download (1MB) | Preview |
Abstract
Multimodal misinformation combining manipulated images with deceptive text poses unprecedented threats to information integrity across social media platforms, affecting over 3.8 billion users globally. Traditional single-modality detection systems achieve only 45-80% accuracy, failing to identify sophisticated cross-modal manipulations that exploit the synergy between visual and textual deception. Current methods process text and images through separate pipelines, missing critical cross-modal inconsistencies and providing no transparent reasoning for detection decisions.
This research proposes an integrated AI framework combining LSTM text classification, GPT-4o multimodal analysis, and Google Fact Check Tools API verification through a six-stage sequential pipeline. The system employs weighted risk scoring: LSTM original text (25%), LSTM generated content (15%), GPT-4o consistency evaluation (25%), GPT-4o detection (25%), and fact-checking (10%). The framework processes textual and visual content through unified multimodal analysis, addressing pipeline separation issues while implementing transparent risk assessment with explainable reasoning for each detection decision.
Evaluation on 1,500 multimodal samples demonstrates 92.0% accuracy [95% CI: 90.6-93.4%], significantly outperforming existing approaches. The system achieves 93.7% precision, 95.0% recall, with operational viability at $0.038 per sample, establishing new benchmarks for multimodal misinformation detection.
Actions (login required)
![]() |
View Item |
Tools
Tools