Data Engineer
Current● Designed and developed end-to-end data pipelines using Azure Databricks, AWS EMR, AWS Glue, and PySpark to process and transform large volumes of data for analytics and reporting purposes.● Implemented data ingestion processes from various sources, including databases, APIs, and flat files, ensuring data integrity and quality.● Leveraged Pandas for data manipulation and cleaning tasks, performing exploratory data analysis to identify trends, anomalies, and insights.● Collaborated with cross-functional teams to gather requirements, understand data needs, and deliver customized solutions to meet business objectives.● Optimized data processing workflows by fine-tuning Spark configurations, partitioning strategies, and utilizing cluster resources efficiently.● Implemented data governance and security measures to ensure compliance with industry standards and regulations.● Integrated AWS services, such as S3, EMR, Lambda, and Athena, to build scalable and cost-effective data processing solutions in the AWS ecosystem.● Utilized Azure services, including Azure Data Factory and Azure Databricks to design and deploy data pipelines in the Azure cloud platform.● Created data models and schemas to support efficient data storage, retrieval, and analysis, ensuring data consistency and accuracy.● Documented data processes, workflows, and system configurations to facilitate knowledge sharing and ensure maintainability.