Data Engineer
Current• Involved in Requirement gathering, business Analysis, Design and Development, testing and implementation of business rules.• Understand business use cases, integration business, write business & technical requirements documents, logic diagrams, process flow charts, and other application related documents.• Used Pandas in Python for Data Cleaning and validating the source data.• Created pipelines, data flows and complex data transformations and manipulations using Glue and PySpark.• Developed ETL applications using Python, Spark (PySpark) and Shell scripting based on the business requirements.• Involved with Data Profiling for multiple sources and answered complex business questions by providing data to business users.• Wrote complex SQL queries for validating the data against different kinds of Database systems to reconcile data across systems such as Snowflake and Oracle.• Developed CI/CD system with Jenkins on Docker for the runtime environment for the CI/CD system to build, test and deploy.• Worked on PySpark jobs and troubleshooting of PySpark job performance. Worked on Batch Data pipelines using PySpark and Scala• Good experience working with Databricks Notebook platform and Snowflake Cloud Data Warehouse systems.• Worked on building centralized Data Lake on AWS Cloud by utilizing services like SQS, SNS, S3, EMR, RedShift and Athena.• Leveraged snowflake’s Snow Sight and Snow SQL to ingest and build data models for downstream analytics.• Worked on Data Ingestion from external systems into S3 data lake using python and boto3 module.• Worked on Docker files to containerize the spark application.