Senior Data Engineer
CurrentCollaborated with cross-functional teams to implement SDLC best practices, enhancing the development and deployment of scalable data pipelines. Optimized workflows with AWS Glue, reducing processing times by 25% through automated metadata management. Designed and deployed pipelines using Apache Spark, Hadoop, and MapReduce, improving data handling capacity by 40%.Integrated AWS RDS and Amazon S3 for secure, scalable data storage, reducing costs by 15%. Developed MongoDB schemas for structured and unstructured data, implementing replication and sharding to ensure high availability. Automated data transformation with Python, Spark SQL, and Pandas, enabling efficient data analysis and preparation.Utilized AWS Kinesis for real-time data streaming and AWS Lambda for cost-effective serverless workflows. Monitored and troubleshot applications using AWS CloudWatch, ensuring system reliability. Streamlined CI/CD processes with AWS CodePipeline and CodeBuild to ensure smooth deployments.Enhanced real-time analytics with interactive visualizations using Presto, D3.js, and Redshift Spectrum, boosting accuracy and engagement by 20%. Optimized queries with Apache Hive, reducing processing times by 35%. Improved collaboration in Databricks, increasing team productivity by 25%.Implemented Docker and Kubernetes for scalable and reliable containerized applications. Established secure data lakes with AWS Lake Formation, enforcing governance and compliance policies. Automated data cataloging for consistent metadata management.Enhanced Amazon Redshift performance by 25% through optimized queries and ensured seamless integration of SQL and NoSQL datasets. Leveraged NumPy and Pandas for advanced data analysis, delivering actionable insights. Ensured compliance with security policies and applied governance best practices to maintain data integrity and traceability.