Senior Aws Data Engineer
CurrentDeveloped scalable and cost-effective AWS Big Data architecture for data lifecycle management including collection, ingestion, storage, processing, and visualization of large, rapidly changing datasets. Implemented end-to-end data pipelines using PySpark for data ingestion into HIVE tables, performing data cleansing, aggregation, and de-duplication. Designed and deployed multi-tier applications using AWS services like EC2, Route53, S3, RDS, DynamoDB, SNS, SQS, IAM with a focus on high availability and fault tolerance.Leveraged Spark features for efficient data preprocessing, scheduling jobs with Airflow, and writing unit tests for Spark code. Managed data extraction from multiple sources, Glue Catalog creation, and Lambda-based automated data transformation. Automated job scheduling with UNIX shell scripts and Cron jobs. Utilized Oozie for data loading into HDFS, Spark streaming with Kafka, and various Spark transformations for data cleansing. Used Jira for issue tracking, Jenkins for CI/CD, and Glue for merging data from multiple tables. Implemented monitors, alarms, and logs using CloudWatch. Assessed AWS services like Amazon EMR, Redshift, and S3 for end-to-end architecture and implementation. Mentored analysts on building analytics tables and collaborated with DevOps for NIFI Pipeline integration with Spark, Kafka, and Postgres.