Data Engineer
Current* Developed high-performance, robust ETL / ELT data pipelines to migrate over 500 million rows of historical data, ensuring data integrity and consistency throughout the process. * Automated manual data operations using Python (Pandas, Great Expectations, PySpark) and SQL, achieving a 70% improvement in operational efficiency and significantly reducing error rates.* Designed and maintained scalable data integration processes with AWS Redshift and Google BigQuery, enabling on-demand business report generation with real-time delivery to clients.* Leveraged Apache Spark to process large retail datasets for ad hoc migration activities for various customers. * Implemented Data Models using Kimball Dimensional Modelling techniques for various data use cases.* Implemented rigorous data quality checks with Great Expectations python library, ensuring that data met predefined standards before integration, leading to enhanced accuracy and reliability of migrated data.