Aws Data Engineer
Current• Developed Spark Applications by using Python and Implemented Apache Spark data processing Project to handle data from various RDBMS and Streaming sources. • Responsible for building scalable distributed data solutions using Hadoop• Extensive experience in working with AWS Cloud Platforms (EC2, S3, EMR, Redshift, Lambda and Glue).• Built NiFi dataflow to consume data from Kafka, make transformations on data, place in HDFS and exposed port to run spark streaming job.• Migrated an existing on-premises application to the AWS platform• Experienced in Maintaining the Hadoop cluster on AWS EMR • Worked with Spark for improving performance and optimization of the existing algorithms in Hadoop.• Work on Spark RDD, Data Frame API, Data set API, Data Source API, Spark SQL, and Spark Streaming.• Used Spark Streaming APIs to perform transformations and actions on the fly for building common.• Developed Kafka consumer API in python for consuming data from Kafka topics. • Consumed Extensible Markup Language (XML) messages using Kafka and processed the XML file using Spark Streaming to capture User Interface (UI) updates.• Created Spark Streaming jobs using Python to read messages from Kafka & download JSON files from AWS S3 buckets.