Sr. Data Engineer
Current• Designed and developed ETL (Extract, Transform, Load) pipelines using AWS services such as Glue, Lambda, and S3 to extract data from various sources, transform it and load it into a data warehouse.• Build and maintain scalable and highly available data solutions using AWS services such as Redshift, RDS, and DynamoDB.• Extensively worked with the application data present in AWS S3 storage for performing required transformations.• Configured and deployed production-ready data… Show more • Designed and developed ETL (Extract, Transform, Load) pipelines using AWS services such as Glue, Lambda, and S3 to extract data from various sources, transform it and load it into a data warehouse.• Build and maintain scalable and highly available data solutions using AWS services such as Redshift, RDS, and DynamoDB.• Extensively worked with the application data present in AWS S3 storage for performing required transformations.• Configured and deployed production-ready data pipelines using AWS Glue, S3 and Lambda to maximize data extraction, transformation, and loading.• Designed and implemented data ingestion pipelines using Apache Airflow to efficiently transfer and process data from various sources.• Created efficient DAGs to represent data workflows, defining task dependencies and execution order.• Integrated Apache Airflow with Amazon S3 to store and retrieve data, ensuring seamless data flow between different stages of the pipeline.• Optimized DAGs for parallel execution, improving data processing performance and reducing processing times.• Performed transformations on the data using Glue jobs with the pyspark script and stored transformed data to S3.• Integrated Apache Airflow with AWS Glue Data Catalog for centralized metadata management and cataloging of data assets within the data pipeline.• Optimized EMR usage and established scalable, fault-tolerant ETL pipelines to facilitate data extraction from multiple sources for storage in the S3 data lake, increasing output efficiency.• Set up CI/CD pipelines to automate testing and deployment of Airflow DAGs, enhancing the development and release process.• Implementing AWS Big Data services such as EMR, Kinesis, and Athena to process large volumes of data.• Utilized AWS Athena interactive querying service for running Ad-hoc queries to perform data analysis on transformed data Show less