Data Engineer
Current• Developing various functionalities for the Data Ingestion framework using Python, resulting in a 30% reduction in data processing time.• Creating the JSON configuration files to cleanse the data through the generic Spark framework.• Developed Spark SQL Jobs and processed data in AWS S3 with complex business requirements for Data Marts creation for various markets.• Developed A/B testing framework to evaluate model improvements, boosting user satisfaction scores by 15%.• Pioneered the integration of AWS Glue for workflow automation, leading to a 25% increase in efficiency and timely data delivery.• Conceptualized and executed complex Spark SQL Jobs, achieving a 40% improvement in processing large-scale data sets for Data Marts.• Successfully implemented SCD2 logic using Spark for master reference data, enhancing historical data accuracy by 15%.• Good exposure in writing SQL queries in data analysis and fixing data issues for various facts and • dimensions for data warehouse and data marts for star schema and snowflake schema.• Managed data migration from Teradata to Snowflake, facilitating improved data analytics capabilities.• Utilized AWS services such as EMR to run Hadoop and Spark. Jobs, S3 for data storage, and CloudWatch for monitoring. Experienced in AWS Lambda, Athena, Glue, and RDS.• Spearheaded the integration of Apache Airflow and ML flow into CI/CD pipelines, automating model versioning and updates, boosting operational efficiency by 30%, and reducing manual errors.