Date of Award
2026
Document Type
Thesis
Degree Name
Master of Science (MS)
Department
Computer Science
Committee Chair
Tathagata Mukherjee
Committee Member
Jacob Hauenstein
Committee Member
Vaidyanath Areyur Shanthakumar
Research Advisor
Tathagata Mukherjee
Subject(s)
Data compression (Computer science), Information retrieval, Artificial intelligence, Neural networks (Computer science)
Abstract
Dense retrieval systems, such as those used in Retrieval Augmented Generation (RAG), keep large embedding vectors in memory, which becomes expensive at scale. This thesis studies how far these embeddings can be compressed without losing retrieval quality. Three encoders (BGE, RoBERTa, and MPNet) are fine-tuned on the same data and evaluated on NanoBEIR using post-hoc methods (INT8, binary, Product Quantization (PQ), TurboQuant), training-time binarization (Straight-Through Estimator (STE), annealed tanh), Matryoshka Representation Learning (MRL), and stacked combinations. No single technique dominates across all backbones. For BGE, which was pre-trained for retrieval, post-hoc PQ leads the Pareto frontier at every tested compression ratio. For RoBERTa and MPNet, training-time binarization is on the frontier instead. Annealed tanh combines with MRL more cleanly than STE, and stacked MRL + PQ remains usable up to 96x compression. Results are presented as Pareto frontiers, so the right method can be selected for a given memory budget and backbone.
Recommended Citation
Awale, Sajil, "Embedding compression for dense retrieval : a controlled comparison of post-hoc, training-time, and stacked methods" (2026). Theses. 853.
https://louis.uah.edu/uah-theses/853