Senior Data Engineer
CurrentUtilized Kafka for the purpose of activity tracking as well as log aggregation.• Skilled in the use of Spark Core, Spark SQL, Spark MLlib, Spark GraphX, and Spark Streaming to process and transform complex data using in-memory computing capabilities built in Scala. I worked with Spark to make existing algorithms more efficient by leveraging Spark Context, Spark SQL, Spark MLlib, Data Frame, Pair RDD's, and Spark YARN.• Experience in designing and implementing complex PL/SQL code to manage healthcare-related data.• Implemented AWS Step functions to automate and orchestrate the Amazon Sagemaker related tasks such as publishing data into S3,training ML model and deploying it.• Designed and implemented serverless ETL pipelines using AWS Glue to automate data ETL processes for healthcare data sources. • Implemented and managed AWS Glue Data Catalog and Created custom classifiers and schema evolution scripts.• Utilized Sqoop export/import to extract the data from Teradata and MySQL and place it in HDFS.• Experience in developing interactive analysis, batch processing, and stream processing applications using PySpark and Spark-Scala is a valuable skill.• Designed and development of data integration solutions using Ab Initio ETL tool to consolidate patient data from sources.• Used Terraform for creating infrastructure as code, managed AWS resources and maintained consistency across environments.