Lead Data Engineer
• Managed a team of 3 data engineers in a banking regulatory project, overseeing requirement gathering, architecture design, and big data application development.• Solely developed a robust, distributed ETL framework using PySpark, adeptly handling thousands of intricate business logics configured in MariaDB, resulting in an impressive 70% reduction in development time.• Solely developed a custom scheduler application in Python to streamline the parallel execution of multiple data… Show more • Managed a team of 3 data engineers in a banking regulatory project, overseeing requirement gathering, architecture design, and big data application development.• Solely developed a robust, distributed ETL framework using PySpark, adeptly handling thousands of intricate business logics configured in MariaDB, resulting in an impressive 70% reduction in development time.• Solely developed a custom scheduler application in Python to streamline the parallel execution of multiple data pipelines across 11 sectors, ensuring the maintenance of dependencies between sectors.• Built a sophisticated data validation tool allowing for user-configured checks, guaranteeing data quality at an impeccable 100% accuracy rate.• Independently developed a data compaction utility in Spark to tackle the small file issue, optimizing big data cluster performance and yielding a substantial 40% enhancement in system performance.• Diagnosed the root cause and bottleneck of slow-running existing big data pipelines, then optimized them by rewriting inefficient code and fine-tuning configurations. This initiative led to a remarkable 60% reduction in the end-to-end data pipeline duration.• Contributed to the development of a data migration pipeline, transferring data from an in-house warehouse (ADA platform) to AWS Redshift for enhanced data analytics and visualization. Additionally, provisioned various AWS services including EC2, SQS, S3, Lambda, API Gateway, and Rekognition for a marketing campaign project on the AWS cloud. Show less