Data Engineer
CurrentCreated connections to various data sources such as amazon redshift, DynamoDB, etc. using spark.Developed spark-based ingestion framework for ingesting data into NoSQL Databases and created tables in Redshift and DynamoDB and executing complex computations and parallel data processing. Worked on Launching various AWS resources such as EC2, S3, Redshift, DynamoDB, etc. using Terraform. Developed Spark code using python for Pyspark and Spark-SQL for faster testing and processing of… Show more Created connections to various data sources such as amazon redshift, DynamoDB, etc. using spark.Developed spark-based ingestion framework for ingesting data into NoSQL Databases and created tables in Redshift and DynamoDB and executing complex computations and parallel data processing. Worked on Launching various AWS resources such as EC2, S3, Redshift, DynamoDB, etc. using Terraform. Developed Spark code using python for Pyspark and Spark-SQL for faster testing and processing of data.Moved the data from RDBMS to HDFS using Sqoop.Scheduled the EMR jobs using AWS Data Pipeline service.Performed Data Validation, Data Cleaning, features scaling, features engineering using Pandas and Numpy packages in Python.Written the code to store the dataflow moment logs into DynamoDB Show less