Lead Data Scientist
CurrentDevelopment: Designed and implemented end-to-end machine learning pipeline in Python to assign risk scores to vehicles using Gradient Boosting Machines. This monthly batch scoring process supports a very large book of business. The model was trained with a 1TB+ dataset in H2o and Spark.◦ Data analysis: Completed a series of ad-hoc analyses to support business operations in a highly regulated industry. Highlights include responding to Department of Insurance inquiries, post-hoc interpretability of machine learning models via SHAP Values, and vendor data evaluation reports to product owners (usually geospatial).◦ Telematics: proprietary work to assess policyholders' risk using sensor data.