Cloud Data Engineer
Current• Leading the architecture and establishment of an Enterprise Data Lake to accommodate diverse use cases, including Analytics, processing, storage, and Reporting of large-scale, dynamically evolving data.• Ensuring the integrity and quality of reference data within the source by executing operations such as cleaning, transformation, and maintenance in a relational environment, collaborating closely with stakeholders and solution architects.• Devising and implementing a Security Framework to grant fine-grained access to objects within AWS S3 utilizing AWS Lambda and DynamoDB.• Designing and managing AWS CDK stacks for deploying serverless applications, ensuring optimal resource allocation.• Leveraging AWS Glue for ETL processes, supporting analytics, processing, storage, and reporting of large-scale, dynamically evolving data.• Utilizing Athena queries for ad-hoc analysis on stored data, facilitating streamlined reporting.• Configuring and operating Kerberos authentication principals to establish secure network communication on the cluster, conducting testing on HDFS, Hive, Pig, and MapReduce for access by new users.• Conducting comprehensive Architecture & implementation assessment of various AWS services such as Amazon EMR, Redshift, and S3.• Regularly conducting performance analysis using CloudWatch metrics to optimize resource utilization and identify areas for enhancement.• Implementing machine learning algorithms using Python to forecast user order quantities for specific items, leveraging Kinesis Firehose and S3 data lake.• Utilizing Spark SQL for Scala & Python interface, automating the conversion of RDD case classes to schema RDD.• Executing data migration from AWS S3 bucket to Snowflake by developing custom read/write snowflake utility functions using Scala.• Importing data from diverse sources such as HDFS/HBase into Spark RDD and conducting computations using PySpark to generate output responses.