Senior Data Engineer
Current• Implemented Partitioning, Dynamic Partitions, Buckets in HIVE.• Working extensively on Hive, SQL, Scala, Spark, and Shell.• Involved in creating ETL pipelines to ingest data into Azure datalake.• Experience in designing data-driven solutions.• Involved in building new data sets, and products helping support business initiatives.• Built production quality ingestion pipelines with automated quality checks to help enable the business to access all the data sets in one place.• Automate the pipeline using Airflow by creating dependencies between each job and schedule them based on daily/weekly/monthly and perform quality checks based on consistency, accuracy, completeness, orderliness, uniqueness, and timeliness and use them to determine ETL job.• Developed Spark core and Spark SQL scripts using Scala for faster data processing. Transformed the data using Spark applications for analytics consumption.• Worked on importing data from SQL Server and Oracle Server into Azure Data Lake using Sqoop.• Worked on creating incremental spark jobs to move data from Azure to Snowflake.• Experience working with delta tables in Azure for loading incremental data into Azure.• Experience in using Azure blob storage for storing various data files.• Developed Managed File Transfer jobs for data exchange to and from third-party vendors and Optum.• Developed Scala script for ingesting flat files, csv, json into Data Lake• Worked on giving production support for developed applications to resolve Critical/Priority-1 issue.• Used Kafka connector for reading CDC feed from Cosmos Db, mongo DB in real time into Kafka• Worked on multiple streaming jobs to read data from Kafka, transform and write into Azure data lake• Support our Data Scientists by helping enhance their modeling jobs to be more scalable when modeling across the entire data set.• Experience in productionize and optimizing data science models.