Sr. Data Engineer
CurrentI specialize in ensuring the reliability and accuracy of data processing within Azure Data Lake by incorporating automated testing frameworks. I've developed incident response plans for quick issue resolution, minimizing downtime in Azure Data Lake and Azure Databricks.For efficient collaboration, I implemented version control for data artifacts and code within Azure Blob Storage. I showcased the capabilities of Azure Stream Analytics by integrating Azure Event Hub and Stream Analytics with Power BI and Azure ML.My expertise extends to developing scalable data processing pipelines using Apache Spark in Scala for efficient ETL operations. I've conducted Proof of Concepts (POCs) comparing Spark and Scala performance with Hive and SQL, deploying them on the Yarn Cluster.Comprehensive testing of Scala applications in big data environments ensures reliability and performance. Performance analysis and optimization strategies are applied to enhance the efficiency of Hadoop clusters.I maintain clear documentation for Python scripts, facilitating collaboration. Python code is developed for workflow management and automation using the Airflow tool. SQL scripts are created for data migration, transformation, and integration tasks within PostgreSQL.I've ensured regulatory compliance for MongoDB databases and managed Relational Database Management Systems (RDBMS) like MySQL and PostgreSQL, optimizing queries and ensuring data integrity.I follow GitLab CI/CD best practices to automate testing, building, and deployment of data engineering solutions, resulting in accelerated development cycles and reduced errors.