Data Engineer
Current• Ensure the on-time delivery of daily files to consumers from managed pipelines and automated workflows on production support utilizing PagerDuty and One ServiceNow incident management• Design and implement CI/CD data pipelines with Jenkins and Git to build and promote applications from QA to Production while meeting necessary testing requirements and staging environments• Build Databricks environments with cluster appropriation, computed resources and performances for East to West transitions to continue team notebook collaboration and scheduled workflows• Promote features and QA processes to live Production environments after testing, verification, development and scheduling in automated engines, including Automic and Redwood servers• Develop transformations using Python, Pyspark, SQL and Scala to build financial processes and modeling for analysis and performance enhancement using collaborative Databricks notebooks and Github repositories• Assist with the migration of older transformations from Java and Scala to current Pyspark environments using Docker and AWS containerization• Manage AWS infrastructure including S3 buckets, versioning, policies, lambdas, EMR and EC2 cluster rehydration for production performance using shell scripting with Unix and Linux in IAM roles.• Register and validate datasets to AWS S3 Data Lakes and Delta Lake environments from QA to live production following stringent data governance, privacy, and data cataloging guidelines using internal applications and platforms• Uphold security standards using vulnerability fixes with Maven, package library updates in Artifactory and ensuring code coverage with compliance standards and testing maturity• Load data from AWS S3 into Snowflake tables, maintain, build and structure tables to optimize data ingestion and improve query performance for financial analysis