Author

Date of Award

2026

Document Type

Thesis

Degree Name

Master of Science (MS)

Department

Computer Science

Committee Chair

Tathagata Mukherjee

Committee Member

Jacob Hauenstein

Committee Member

Vaidyanath Areyur Shanthakumar

Research Advisor

Tathagata Mukherjee

Subject(s)

Data compression (Computer science), Information retrieval, Artificial intelligence, Neural networks (Computer science)

Abstract

Dense retrieval systems, such as those used in Retrieval Augmented Generation (RAG), keep large embedding vectors in memory, which becomes expensive at scale. This thesis studies how far these embeddings can be compressed without losing retrieval quality. Three encoders (BGE, RoBERTa, and MPNet) are fine-tuned on the same data and evaluated on NanoBEIR using post-hoc methods (INT8, binary, Product Quantization (PQ), TurboQuant), training-time binarization (Straight-Through Estimator (STE), annealed tanh), Matryoshka Representation Learning (MRL), and stacked combinations. No single technique dominates across all backbones. For BGE, which was pre-trained for retrieval, post-hoc PQ leads the Pareto frontier at every tested compression ratio. For RoBERTa and MPNet, training-time binarization is on the frontier instead. Annealed tanh combines with MRL more cleanly than STE, and stacked MRL + PQ remains usable up to 96x compression. Results are presented as Pareto frontiers, so the right method can be selected for a given memory budget and backbone.

Available for download on Thursday, August 05, 2027

Share

COinS