Banala, Saketh Reddy (2025) Deep learning chess framework to synthesize language model reasoning with traditional engines. Masters thesis, Dublin, National College of Ireland.
Preview |
PDF (Master of Science)
Download (1MB) | Preview |
Preview |
PDF (Configuration Manual)
Download (1MB) | Preview |
Abstract
This research investigates comparative large language models' (LLMs) performance against classical chess engine capability in strategic analysis and accuracy of moves across a range of chess positions. While classical engine engines are accurate in calculation, they are devoid of explanation to provide insight to human understanding. Conversely, modern language models are demonstrated to have been capable of strategic reasoning but, to the present understanding, are rarely challenged through formal analysis with a chess scenario. This research work corrects this by constructing a comprehensive evaluation protocol between Gemini 1.5 Pro and Claude 3.5 Sonnet over classical engines Berserk 13 and Caissa 1.22, with Stockfish 17 posing as evaluation arbitrator. The protocol uses a systematic evaluation of 180 games played with 12 engine pairs and 15 positions that were carefully chosen to show different types of scenarios, such as opening, tactics, position, endgame, and strategy. The performance metrics include Centipawn Loss (CPL), Game Point Loss (GPL), and move quality categorisation. These results show that traditional engines are much better than language models, with an average performance gap of 237.24 CPL, or almost 2.37 pawns. Position-specific analysis shows that performance differences are not always the same: 90 CPL during opening positions, 247 CPL during tactical positions, and 109 CPL during positional play. Language models were good at recognising patterns in tasks, but they weren't as good at making accurate tactical calculations. The findings indicate possible advantages of hybrid systems that integrate the computational precision of traditional engines with the pattern recognition and explanatory functions of language models. This research offers empirical evidence of the synergistic strengths inherent in both methodologies and establishes a basis for the creation of an integrated chess analysis system that bridges the divide between machine accuracy and human comprehension.
Actions (login required)
![]() |
View Item |
Tools
Tools