Date of Award

2026

Document Type

Thesis

Degree Name

Master of Science (MS)

Department

Computer Science

Committee Chair

Tathagata Mukherjee

Committee Member

Jacob Hauenstein

Committee Member

Vaidyanath Areyur Shanthakumar

Research Advisor

Tathagata Mukherjee

Subject(s)

Natural language processing (Computer Science), Artificial intelligence, Big data

Abstract

NASA’s GHG Center and VEDA platforms publish satellite Earth observation datasets through STAC catalogs. Publishing a new dataset requires a subject-matter expert to author the discovery configuration and catalog collection for the data, a task that takes three to five hours per dataset to complete. This work contributes a metadata curation pipeline that turns granules, the documentation, and usage examples into both files.Fields readable directly from the files, such as spatial and temporal extent, are extracted deterministically. Fields that require language understanding are sent to a locally hosted small language model. The small model infers the structure of the asset and variable and uses retrieval-augmented generation to write accurate titles and descriptions based on the documentation of the field. The small language model rivals the performance of frontier models and beats them in grounding at zero marginal cost, producing a valid catalog every time. This shifts the subject-matter expert from authoring to review.

Share

COinS
 
 

To view the content in your browser, please download Adobe Reader or, alternately,
you may Download the file to your hard drive.

NOTE: The latest versions of Adobe Reader do not support viewing PDF files within Firefox on Mac OS and if you are using a modern (Intel) Mac, there is no official plugin for viewing PDF files within the browser window.