Senior Data Engineer
CurrentDesign, develop and maintain non-production and production transformations in AWS environment and create the data pipelines using PySpark Programming. Wrote, compiled, and executed programs as necessary using Apache Spark in Scala to perform ETL jobs with ingested data. Worked on Big data on AWS cloud services i.e., EC2, S3, EMR and DynamoDB• Implemented End to End solution for hosting the web application on AWS cloud with integration to S3 buckets. Used Spark Streaming to divide streaming data into batches as an input to Spark engine for batch processing. Involved in designing and deploying multi-tier applications using all the AWS services like (EC2, Route53, S3, RDS, Dynamo DB, SNS, SQS, IAM) focusing on high-availability, fault tolerance, and auto-scaling in AWS Cloud Formation• Wrote Spark applications for data validation, cleansing, transformation, and custom aggregation and used Spark engine, Spark SQL for data analysis and provided to the data scientists for further analysis