Azure Data Engineer
CurrentConducted performance tuning and optimization for Kubernetes and Docker deployments, enhancing overall system performance.Functioned as Data Engineer responsible for data modelling, data migration, design, preparing ETL pipelines for both cloud and onperm.Good knowledge on Apache Spark components including SPARK CORE, SPARK SQL, SPARK STREAMINGDeveloped Spark applications for data validation, cleansing, transformation and aggregation.Configured Spark Streaming to receive real time data from the Apache Kafka and store the stream data to HDFS using Scala and PySpark.Developed MapReduce code for processing and parsing the data from various sources later stored parsed data into HBase, Hive using HBase-Hive Integration. Involved in loading, transforming large sets of Structured, Semi-Structured and Unstructured data later analyzed them by running Hive queriesImplemented Synapse Integration with Azure Databricks, halving development efforts and improving performance through dynamic partition switch implementation.Managed real-time data ingestion using Kafka and Spark Streaming, implementing robust data quality checks for enhanced reliability.Contributed to all phases of the Software Development Lifecycle (SDLC), from requirements gathering to deployment and analysis.Proficient in Hadoop, SOLR, PySpark, Kafka, Storm, and webMethods for big data integration and analytics solutions.