Sr. Data Engineer
Current Designed and developed ETL (Extract, Transform, Load) pipelines usingAWS services such as Glue, Lambda, and S3 to extract data from varioussources, transform it and load it into a data warehouse. Build and maintain scalable and highly available data solutions using AWSservices such as Redshift, RDS, and DynamoDB. Extensively worked with the application data present in AWS S3 storage forperforming required transformations. Experience wif big data processing tools like CouchDB that allows running asingle logical database server on any number of servers. Used DataStax Cassandra along wif Pentaho for reporting. Implemented various Data Modeling techniques for Cassandra Improved teh query performance by transitioning log storage fromCassandra to Azure SQL Datawarehouse. Integrated on-premises data(MySQL, Cassandra) Configured and deployed production-ready data pipelines using AWS Glue,S3 and Lambda to maximize data extraction, transformation, and loading. Designed and implemented data ingestion pipelines using Apache Airflow toefficiently transfer and process data from various sources. Created efficient DAGs to represent data workflows, defining taskdependencies and execution order. Integrated Apache Airflow with Amazon S3 to store and retrieve data,ensuring seamless data flow between different stages of the pipeline. Optimized DAGs for parallel execution, improving data processingperformance and reducing processing times. Performed transformations on the data using Glue jobs with the pysparkscript and stored transformed data to S3. Integrated Apache Airflow with AWS Glue Data Catalog for centralizedmetadata management and cataloging of data assets within the data pipeline. Set up CI/CD pipelines to automate testing and deployment of Airflow DAGs,enhancing the development and release process. Utilized AWS Athena interactive querying service for running Ad-hoc queriesto perform data analysis on transformed data