Summer Research Student
Conducted feasibility studies on automating metadata extraction from research papers to establish metadata standardsUsed Python to download xml documents of research papers, then passed the Methods section to GPT API to scrape metadata terms, using key prompt engineering principlesUsed GPT embedding model to embed the metadata terms, storing terms and embeddings in an SQL databaseClustered the metadata terms using the scikit-learn library to group similar terms and determine the most important and commonly occurring terms