Date of Award
2026
Document Type
Thesis
Degree Name
Master of Science (MS)
Department
Computer Science
Committee Chair
Tathagata Mukherjee
Committee Member
Jacob Hauenstein
Committee Member
Vaidyanath Areyur Shanthakumar
Research Advisor
Tathagata Mukherjee
Subject(s)
Natural language processing (Computer Science), Artificial intelligence, Big data
Abstract
NASA’s GHG Center and VEDA platforms publish satellite Earth observation datasets through STAC catalogs. Publishing a new dataset requires a subject-matter expert to author the discovery configuration and catalog collection for the data, a task that takes three to five hours per dataset to complete. This work contributes a metadata curation pipeline that turns granules, the documentation, and usage examples into both files.Fields readable directly from the files, such as spatial and temporal extent, are extracted deterministically. Fields that require language understanding are sent to a locally hosted small language model. The small model infers the structure of the asset and variable and uses retrieval-augmented generation to write accurate titles and descriptions based on the documentation of the field. The small language model rivals the performance of frontier models and beats them in grounding at zero marginal cost, producing a valid catalog every time. This shifts the subject-matter expert from authoring to review.
Recommended Citation
Rimal, Ram Sharan, "Metadata extraction from satellite data using small language models" (2026). Theses. 842.
https://louis.uah.edu/uah-theses/842