Senior Big Data Engineer
Current• Used Agile methodology in developing the application, which included iterative application development, weekly Sprints, stand up meetings and customer reporting backlogs.• Development of efficient pig and hive scripts with joins on datasets using various techniques.• Automating the ETL tasks and data work flows for the data pipeline of the ingest process through UC4 scheduling tool.• Solved performance issues in Hive and Pig scripts with understanding of Joins, Group and aggregation and how does it translate to MapReduce jobs.• Created data sharing between two snowflake accounts.• Wrote ETL jobs to read from web APIs using REST and HTTP calls and loaded into HDFS using java and Talend.• Created internal and external stage and transformed data during load.• Monitoring resources and Applications using AWS Cloud Watch, including creating alarms to monitor metrics such as EBS, EC2, ELB, RDS, S3, SNS and configured notifications for the alarms generated based on events defined.• Analyze and develop programs by considering the extract logic and the data load type using Hadoop ingest processes using relevant tools such as Sqoop, Spark, Scala, Kafka, Unix shell scripts and others• Assist with the analysis of data used for the tableau reports and creation of dashboards. • Design and implement large scale distributed solutions in AWS.• Optimized Map Reduce Jobs to use HDFS efficiently by using various compression mechanisms.• Used ORC and Parquet file formats in Hive.• Writing code and creating hive jobs to parse the logs and structure them in tabular format to facilitate effective querying on the log data.• Created and managed cloud VMs with AWS EC2 Command line clients and AWS management console.• Migrated on premise database structure to Confidential Redshift data warehouse. Worked on AWS Data Pipeline to configure data loads from S3 into Redshift.