Data Engineer
CurrentIntegrated Azure HDInsight with Azure Data Lake Storage, ensuring seamless data flow and storage for analytics workloads, and enabling comprehensive data exploration and analysis. Utilized advanced data partitioning and indexing techniques in Azure Data Factory and Azure Synapse Analytics to improve query performance and reduce latency. Used Python to clean and pre-process raw data, ensuring data quality and consistency for analysis. Administered and managed the Hadoop ecosystem, including HDFS, Map Reduce, Hive, Pig and Spark, ensuring optimal performance and reliability.Implemented and maintained metadata catalogs for Spark SQL, ensuring data lineage and governance. Worked on Kafka topic partitioning strategies to optimize data distribution and parallel processing. Implemented data governance practices within Kafka and Scala, incorporating metadata management techniques to enhance data discoverability, lineage tracking, and overall data governance. Integrated Scala applications with monitoring and logging tools to proactively identify issues, analyze performance, and streamline troubleshooting processes Participated in the evaluation and implementation of Snowflake features and enhancements to improve data warehousing and analytics capabilities.Managed version control for Tableau workbooks and data sources, facilitating collaboration among team members and providing a clear audit trail for changes. Implemented global filters in Tableau to allow users to dynamically control multiple visualizations simultaneously, enhancing the overall user experience• Experienced in documentation of Power BI solutions, including data models, transformations, report specifications, to facilitate knowledge transfer and future maintenance. Maintained clear and comprehensive project documentation on GitHub for improved project understanding and on boarding. Used Azure DevOps and Jenkins pipelines to build and deploy different resources like Code and Infrastructure in Azure.