Principal Data Engineer
Authored a collection of microservices to support cloud migration from SQL Server data warehouse to AWS data lake. Including common ETL and data engineering workflows: SQLServer-to-Parquet (via PySpark script in EMR), Redshift-to-Redshift, CSV to Redshift.Pipelines/workflows built leveraging AWS serverless architecture including Elastic Map Reduce (EMR), Elastic Container Service (ECS), API Gateway, Lambda, DynamoDB, Redshift, S3. The asynchronous queues are serviced by python-based workflow where Docker images run in ECS tasks within FARGATE containers on an ECS cluster. Used Django framework to present dashboard that tracks the progress of requests within a single pane of glass. Additional configuration functionality offered includes load balancing, throttling, & priority setting.