NORMA eResearch @NCI Library

Dynamic Pricing for Food Waste Reduction in Supermarkets Using Reinforcement Learning

Kataradahalli Sumanth, Sathwik (2025) Dynamic Pricing for Food Waste Reduction in Supermarkets Using Reinforcement Learning. Masters thesis, Dublin, National College of Ireland.

[thumbnail of Master of Science]
Preview
PDF (Master of Science)
Download (586kB) | Preview
[thumbnail of Configuration Manual]
Preview
PDF (Configuration Manual)
Download (381kB) | Preview

Abstract

Supermarkets usually use a fixed price or a plain markdown scheme to sell perishable goods, and thus they lose either too much or too little money. Reinforcement learning (RL) is a data-driven alternative that can adjust the prices according to the demand and inventory changes. This paper creates a simple and reproducible simulator that simulates FEFO stock ageing, stochastic demand and discrete price Multipliers, and determines whether a policy-gradient RL approach, Proximal Policy Optimization (PPO), can learn to perform better than heuristic knowledge of the supermarket. Three policies (Fixed, Random and PPO) are evaluated on a series of simulated 589-day episodes with a reward function which specifically punishes expired stock. On five randomly chosen seeds, PPO will always get to zero waste and competitive revenue with the greatest total payoff despite the fact that it has a conservative operation as compared to the Fixed and Random policies depending on the waste penalty. The Random policy has erratic performance and average waste whereas the Fixed baseline maximizes revenue but results in extreme waste which is operationally impossible. These results indicate that even a naive RL policy can encode perishability constraints and be more helpful in providing a better revenue-waste tradeoff than the usual retail heuristics. The future work must ensure that it has more demand modelling, continuous pricing and operations constraints to ensure that it can be deployed in real supermarket environments.

Item Type: Thesis (Masters)
Supervisors:
Name
Email
Nolan, Eamon
UNSPECIFIED
Uncontrolled Keywords: Reinforcement learning; Proximal Policy Optimization (PPO); dynamic pricing; perishable goods; retail supermarket; FEFO inventory ageing; waste reduction; sustainability; revenue management; stochastic demand; Gymnasium simulator; stable-baselines3
Subjects: Q Science > QA Mathematics > Electronic computers. Computer science
T Technology > T Technology (General) > Information Technology > Electronic computers. Computer science
T Technology > TD Environmental technology. Sanitary engineering
H Social Sciences > HD Industries. Land use. Labor > Specific Industries > Food Industry
H Social Sciences > HD Industries. Land use. Labor > Specific Industries > Retail Industry
Divisions: School of Computing > Master of Science in Data Analytics
Depositing User: Ciara O'Brien
Date Deposited: 07 Sep 2026 13:30
Last Modified: 07 Sep 2026 13:30
URI: https://norma.ncirl.ie/id/eprint/9869

Actions (login required)

View Item View Item