Senior Data Engineer
Current• Worked with Hadoop ecosystem and Implemented Spark using Python and utilized Data frames and Spark SQL API for faster processing of data.• Developed Spark Streaming job to consume the data from the Kafka topic of different source systems and push the data into HDFS locations.• Data sources are extracted, transformed, and loaded to generate CSV data files with Python programming and SQL queries.• Developed Automation Regressing Scripts for validation of ETL process between multiple databases like AWS Redshift, Postgres, MongoDB, T-SQL, and SQL Server using Python. • Used Airflow for scheduling the Hive, Spark, and MapReduce jobs.• Developed and managed ETL workflows on Snowflake using AWS infrastructure, automating data integration from diverse sources while ensuring data integrity and security through IAM roles and encryption mechanisms.• Converting Hive/SQL queries into Spark transformations using Spark RDDs and Pyspark• Filtering and cleaning data using Python code and SQL Queries• Created Snow pipe for continuous data load from staged data residing on cloud gateway servers.• Troubleshooting errors in Hbase Shell/API, Pig, Hive and MapReduce.• Implemented Installation and configuration of multi-node cluster on Cloud using Amazon Web Services (AWS) on EC2.