Senior Data Engineer
CurrentI have extensive experience in designing and implementing scalable data pipelines using Azure Data Factory and developing optimized data models in Azure Synapse Analytics to support complex queries and analytics. My work includes managing data storage solutions such as Azure Data Lake Storage and Azure Blob Storage, as well as designing ETL processes and creating high-level design documents for data flow, extraction, and staging. I have a solid background in developing and optimizing Informatica mappings, orchestrating data pipelines, and using tools like SQL Server for creating complex SQL queries, views, and managing data in dimensional and relational data warehouses.In my role, I have developed Spark applications in both PySpark and Scala to handle data cleansing, transformation, and ingestion for machine learning and reporting purposes. I have implemented workflows in Hadoop using MapReduce, optimized HDFS storage, and scheduled automated data processing jobs using Control-M. My experience also extends to implementing and maintaining machine learning models in production environments with TensorFlow and PyTorch, and developing interactive Tableau dashboards and reports to provide actionable business insights.Additionally, I have managed continuous integration workflows using GitHub Actions and integrated Azure DevOps with Azure Repos for version control and automated code reviews. I have also maintained repositories on GitLab, managed user permissions, and created Jira tickets for effective team communication on data pipeline issues and enhancements. My proactive involvement in status meetings and updates has ensured seamless project management and alignment with team goals.