Senior Azure Data Engineer
CurrentContributed to the development of PySpark Data Frames in Azure Data bricks to read data from Data Lake or Blob storage and utilize Spark SQL context for transformation.Developed ETL pipelines in and out of data warehouse using a combination of Python and Snowflake. Used Snow SQL to write SQL queries against Snowflake.Worked on an Azure copy to load data from an on-premises SQL server to an Azure SQL Data warehouse.Integration of data storage solutions in spark - especially with Azure Data Lake storage and Blob snowflake storage.Developed data ingestion pipelines on Azure HDInsight Spark cluster using Azure Data Factory and Spark SQL.Designed and implemented a real-time data streaming solution using Azure EventHub.Developed Spark Streaming applications to process real-time data from various sources, such as Kafka and Azure Event Hubs.Built streaming ETL pipelines using Spark Streaming to extract data from various sources, transform it in real- time, and load it into a data warehouse like Azure Synapse Analytics.Developed Spark APIs to import data into HDFS from Teradata and created Hive tables.Loaded data into Parquet Hive tables from Avro Hive tables.Demonstrated proficiency in programming languages like Python and Scala.Performed various types of joins on Hive tables and implemented Hive SerDe like Avro, JSON.Developed Spark core and Spark SQL scripts using Scala for faster data processing.Worked on the Hadoop ecosystem in PySpark on HDInsight and Databricks.Created reports in Power BI with reference to wire frame documentation.Implemented DAX expressions for MTD and YTD based on slicer selection.Configuration of RLS (Row Level Security) for specified users with the help of specific user selection in Power BI service.Configuration of Power BI service account after publishing reports Extensively utilized Spark core, Spark SQL, and Spark Streaming for real-time data processing.Conducted data profiling and transformation on raw data using Python.