Venkatesan, Keerthana (2025) Enhancing Legal Research Efficiency through Retrieval-Augmented Generation (RAG) with Vector Search. Masters thesis, Dublin, National College of Ireland.
Preview |
PDF (Master of Science)
Download (634kB) | Preview |
Preview |
PDF (Configuration Manual)
Download (697kB) | Preview |
Abstract
The legal retrieval of information has unique problems because of the complex legal language used, citation formats, and requirements of specific output with even explanations. In this paper, discussed the creation of a legal chatbot assistant based on Retrieval-Augmented Generation (RAG) together with domain-oriented large language models that respond to legal questions with references to documents uploaded in the process. To achieve such a contextual search, PDF files are semantically divided into chunks and indexed with the help of vector embeddings in a Qdrant vector database and stored. There is text retrieval and resulting grounded responses in real-time to the queries. Also, evaluate different subsystems in the pipeline, like the effectiveness of text chunking, trade-offs in optimising embedding, and latency as far as document size is concerned. An efficient vector representation was provided by sentence-transformers/all-MiniLML6- v2, which is a fine-tuning of the embedding model that resulted in high retrieval recall. Latency tests were scalable, with less than 1.2 seconds obtained as the average response time on querying a document library with a size of 100+ documents. To test the performance of the system, various metrics, which included ROUGE-L, BLEU, latency, and retrieval hit rate, were used on a benchmark question-answering dataset involving legal questions. It proves that it was accurate and relevant. Visualization was used to read the performance of the system. This system shows the possibility of the usage of AI-driven means of legal support that guarantees quicker, verifiable, and easily accessible legal support.
Actions (login required)
![]() |
View Item |
Tools
Tools