Mallik, Roopak (2025) A Comparative Study of Reinforcement Learning Algorithms and Custom Technical Indicators for Short-Term Stock Trading. Masters thesis, Dublin, National College of Ireland.
Preview |
PDF (Master of Science)
Download (4MB) | Preview |
Preview |
PDF (Configuration Manual)
Download (2MB) | Preview |
Abstract
Reinforcement learning (RL) has emerged as a promising alternative to traditional algorithmic trading approaches, yet significant gaps remain in understanding how different RL algorithms compare and whether adding technical indicators improves trading performance. This dissertation systematically compares three prominent RL algorithms—Deep Q-Network (DQN), Proximal Policy Optimization (PPO), and Advantage Actor-Critic (A2C) for short-term stock trading using historical daily data from three GPU technology companies (NVIDIA, AMD, and Intel) over eight months from January to August 2025.
The research addresses two key objectives: first, comparing DQN, PPO and A2C performance using standard market data (Open, High, Low, Close, Volume); second, determining whether adding technical indicators (RSI, SMA, OBV) improves trading outcomes. Using OpenAI Gymnasium, agents were trained on 20,000 timesteps of historical data and evaluated on separate test data under identical conditions to ensure fair comparison. Performance was measured using factors such as total profit.
Results show that A2C outperforms the other algorithms, occasionally achieving profits above break-even (1.0385 on AMD with standard data), though with significant variability between runs. DQN generally remains near break-even but experiences substantial losses on Intel, while PPO consistently underperforms, never reaching profitability. Adding technical indicators produces relatively moderate improvements across all algorithms, with A2C benefiting most (reaching 1.0652 on AMD and 1.0749 on Intel). However, no algorithm achieves consistent profitability, highlighting fundamental challenges in applying RL to volatile financial markets with noisy data.
This research advances the field by establishing a standardized experimental framework that enables fair algorithm comparison and demonstrates that while algorithmic choices and feature engineering influence performance, significant challenges remain in developing stable and consistently profitable automated trading systems. The findings provide realistic expectations for both researchers and practitioners, emphasizing the need for rigorous evaluation, careful feature engineering and acknowledgment of limitations when deploying RL-based trading strategies in real-world financial markets.
| Item Type: | Thesis (Masters) |
|---|---|
| Supervisors: | Name Email Hamill, David UNSPECIFIED |
| Subjects: | H Social Sciences > HG Finance Q Science > QH Natural history > QH301 Biology > Methods of research. Technique. Experimental biology > Data processing. Bioinformatics > Artificial intelligence Q Science > Q Science (General) > Self-organizing systems. Conscious automata > Artificial intelligence Q Science > Q Science (General) > Self-organizing systems. Conscious automata > Machine learning H Social Sciences > HG Finance > Investment > Stock Exchange |
| Divisions: | School of Computing > Master of Science in Data Analytics |
| Depositing User: | Ciara O'Brien |
| Date Deposited: | 08 Sep 2026 09:06 |
| Last Modified: | 08 Sep 2026 09:06 |
| URI: | https://norma.ncirl.ie/id/eprint/9877 |
Actions (login required)
![]() |
View Item |
Tools
Tools