Data Engineer
Current• Developed and deployed scalable data solutions on AWS and Azure platforms, leveraging services like AWS Glue, Azure Data Factory, and Azure Data Lake Storage.• Led data migration projects from on-premises systems to cloud platforms, ensuring seamless transition and optimizing performance.• Reduced data migration time by 25% through improved data ingestion and transformation techniques using AWS Glue and Azure Data Factory.• Utilized cloud-native features for data storage, processing, and analytics, working with AWS Redshift, Azure SQL, and AWS S3.• Optimized queries and managed metadata using Hive for querying and analyzing large datasets in distributed storage environments like Hadoop.• Executed large-scale data processing and analytics with Spark, employing Spark SQL, Spark Streaming, and Spark for various data manipulation and analysis tasks.• Optimized data processing time by 30% by implementing efficient Spark and Hive queries.• Designed and developed data analysis solutions using Scala within Hadoop ecosystems like Spark, handling complex data transformations and computations effectively.• Engineered and optimized ETL processes for data extraction, transformation, and loading from various sources using tools like AWS Glue, Apache Spark, and custom Python scripts.• Implemented batch and stream processing techniques for efficient data ingestion and processing, utilizing AWS Kinesis, AWS Data Pipeline, and Apache Kafka.• Performed data manipulation, loading, extraction, and analysis using Python, leveraging libraries like Pandas, NumPy, and SciPy.• Wrote complex SQL queries and stored procedures to query and manipulate data in relational databases like PostgreSQL, MySQL, and SQL Server, optimizing database performance.• Designed and implemented data warehousing solutions for storing and analyzing structured data efficiently.