Date of Award
2026
Document Type
Thesis
Degree Name
Master of Science (MS)
Department
Computer Science
Committee Chair
Tathagata Mukherjee
Committee Member
Letha Hughes Etzkorn
Committee Member
Jacob Hauenstein
Research Advisor
Tathagata Mukherjee
Subject(s)
Artificial intelligence--Data processing, Natural language processing (Computer science), Computer programs--Correctness, Counterfactuals (Logic)
Abstract
Correctness and faithfulness are standard metrics used for evaluating Retrieval-Augmented Generation (RAG) systems, but they only describe the outcomes and do not surface the underlying behaviors behind them. The same scores can hide a number of underlying behaviors. They can hide whether a system relied on retrieval, answered from parametric knowledge, followed a near-miss context, missed the gold source or failed to use the retrieved evidence. The same scores can hide all of these. This thesis introduces a counterfactual behavioral evaluation framework that tests each query under five controlled context conditions and uses a gold-source coverage gate to identify whether the annotated source passage was retrieved. These observations together are used to assign each query to one of the ten behavioral labels defined in this thesis, which indicate to a RAG practitioner where engineering attention should be directed. The framework has been evaluated across three diverse knowledge domains and two model capability classes. The framework reveals distinctions that are hidden by correctness and faithfulness alone and it treats these metrics as starting points for RAG diagnosis and not as endpoints.
Recommended Citation
Khan, Saif Raza, "What correctness and faithfulness cannot tell you : counterfactual behavioral diagnosis of RAG systems" (2026). Theses. 843.
https://louis.uah.edu/uah-theses/843