Sr. Data Engineer
CurrentInvolved in requirements gathering on the data subjects and analysis of the source data before designing.Build Hadoop Datalakes/Deltalake and developed the architecture and used in implementations within the organization and Ingest Legacy datasets into HDFS using Sqoop Scripts and populate Enterprise DataLake by importing tables from Oracle and other databases and Mainframe Sources and store them in partitioned hive tables using different file compression techniques. Installed and Setup Hadoop CDH clusters for development and production environment and installed and configured Hive, Pig, Sqoop, Flume, Cloudera manager and Oozie on the Hadoop cluster.Created data pipeline of gathering, cleaning and optimizing data using Hive, Spark and Building data pipeline ETLs for data movement to S3, then to Redshift and Developed automated data pipelines from various external data sources (web pages, API etc) to internal data warehouse (SQL server, AWS), then export to reporting tools.Implement AWS Data Lake leveraging S3, terraform, EC2, Lambda, and IAM in performing data processing and storage while writing complex SQL queries, analytical and aggregate functions on views in Snowflake data warehouse to develop near real time visualization using Tableau.Used Spark Streaming APIs to perform transformations and actions on the fly for building common learner data model which gets the data from Kafka in near real time and persist it to Cassandra.Responsible for writing Hive Queries for analyzing data in Hive warehouse using Hive Query Language (HQL) and Hive UDF's in Python.Involved in installing EMR clusters on AWS and used AWS Data Pipeline to schedule an Amazon EMR cluster to clean and process web server logs stored in Amazon S3 bucket and Created monitors, alarms, notifications and logs for Lambda functions, Glue Jobs, EC2 hosts using Cloudwatch