NORMA eResearch @NCI Library

Efficient and Scalable Retail Demand Forecasting with Sparse Attention Mechanisms

Thangellapally, Akhilesh (2025) Efficient and Scalable Retail Demand Forecasting with Sparse Attention Mechanisms. Masters thesis, Dublin, National College of Ireland.

[thumbnail of Master of Science]
Preview
PDF (Master of Science)
Download (705kB) | Preview
[thumbnail of Configuration Manual]
Preview
PDF (Configuration Manual)
Download (507kB) | Preview

Abstract

Accurate retail demand forecasting is critical for inventory management, pricing optimization, and strategic decision-making, yet existing machine learning and deep learning models often face limitations in capturing complex temporal patterns while maintaining computational efficiency. High-dimensional retail datasets containing mixed categorical and numerical attributes place substantial strain on traditional forecasting architectures, particularly transformer-based models with dense attention mechanisms that incur high memory requirements and extended training times. This study addresses these challenges by developing a Sparse Attention–based transformer architecture designed to enhance forecasting accuracy while reducing computational overhead. The research follows a structured methodology involving data preprocessing, exploratory analysis, and the implementation of multiple forecasting baselines, including ensemble machine learning models, neural networks, and recurrent architectures. Categorical features are encoded through embeddings, numerical variables are normalized, and both are integrated within a stack of sparse transformer blocks that selectively attend to relevant neighbourhoods rather than processing all feature interactions uniformly. The models are evaluated using standard regression metrics—MAE, MSE, RMSE, and R²—to provide a comprehensive assessment of forecasting performance. Experimental results demonstrate that the Sparse Attention Model achieves a substantial improvement over all baselines, attaining exceptionally low error values and an R² of 0.9890, indicating strong generalization and highly accurate demand estimation. These findings show that incorporating sparsity into attention mechanisms significantly improves computational scalability without compromising representational capacity. The study concludes that sparse attention architectures offer a robust and efficient alternative for large-scale retail forecasting and provide a promising direction for future research in attention-based modelling.

Item Type: Thesis (Masters)
Supervisors:
Name
Email
Agarwal, Bharat
UNSPECIFIED
Subjects: Q Science > Q Science (General) > Self-organizing systems. Conscious automata > Machine learning
H Social Sciences > HD Industries. Land use. Labor > Specific Industries > Retail Industry
Divisions: School of Computing > Master of Science in Data Analytics
Depositing User: Ciara O'Brien
Date Deposited: 09 Sep 2026 10:27
Last Modified: 09 Sep 2026 10:27
URI: https://norma.ncirl.ie/id/eprint/9920

Actions (login required)

View Item View Item