Head Of Machine Learning
CurrentI lead machine learning and NLP efforts at Statt - working extensively with a large dataset of millions of public policy-related documents. Details include: - Built/trained model to classify documents as relevant to public policy - goal was to weed out bad documents from our data collection. - Built/trained model to classify documents with respect to type (white paper, commentary, press release, etc) - which greatly expanded search capabilities of the product.- Employed over a dozen topic classification models to classify documents as discussing various public policy-related topics such as economics, housing, and transportation.- Developed data collection and labeling strategies for these models for model training from our existing corpus of documents (which number in the millions).- Developed and deployed state-of-the-art entity linking solution in PyTorch which greatly expanded the data collected from each of our documents.- Developed text summarization service built on existing model that was fine tuned with relevant training data. These summaries appear in the application to provide users with a brief synopsis of each article.