Senior Data Engineer
Current● Designed ELT pipelines with 5 stages of data lineage and modeling that are orchestrated by Prefect. Source data is extracted intoAWS S3, after which it is cataloged in AWS Glue and prepared for loading into Redshift. Data is modeled in dbt and exposed toLooker, where it is used to empower company-wide data literacy and analytics. Data pipelines are subject to strict review process tomake sure data adheres to standards of healthcare data privacy (HIPAA and PHI restrictions).● Revamped the Data Engineering team’s CICD process using custom Github actions. Proposed changes to production code aretested in AWS ECR environments to ensure new changes pass necessary validation before going live.● Spearheaded project to design Salesforce data pipelines that enabled cross functional teams to access Salesforce data, as well asreverse ETL processes that enriched Salesforce with data from various external sources and helped executives with planningsurrounding new customers/healthcare provider partners, prospective market expansion, and company budgets.● Launched new data observability tool called Elementary to monitor data model run times and test results. The tool allowedengineers to identify which data models caused bottlenecks and to debug data validity test failures.● Partnered with Data Scientists and ML Engineers to define feature sets to be used for ML models and experimentation. Data wastransformed before being ingested to AWS using the FeatureStore SDK. Models were trained on the datasets using Sagemaker.