Hadoop Developer
Current• Expertise in designing and deployment of Hadoop cluster and different Big Data analytic tools including Pig, Hive, HBase, Oozie, Zookeeper, SQOOP, flume, Spark, Impala, Cassandra with Cloudera Distribution.• Installed Hadoop, Map Reduce, HDFS, AWS and developed multiple MapReduce jobs in PIG and Hive for data cleaning and pre-processing.• Implemented Partitioning, Dynamic Partitions, Bucketing in HIVE.• Created Hive tables with some partition column and stored in ORC format.• Created External tables in hive and stored the data.• Prepared Hive DDLs, queries and successfully executed on top of Spark.• Able to ingest the data onto HDFS and Hive from various platforms like Flat files and Mainframe files using Syncsort.• Created Syncsort task (dxt files) and also jobs (dxj files) with lot of conversions on the given fields.• Performed query optimization techniques using map side joins in hive. • Loading CSV files into data frame and performed Spark RDD transformations, actions to implement business analysis.• Involved in converting Hive/SQL queries into Spark transformations using Spark RDDs.• Prepared Scala jar to process the data on top of Spark.• Involved in deploying the code through Source tree from Git to preprod, from preprod to prod.• Handle multiple issues during batch support.• Experience in importing and exporting data using Sqoop from HDFS to Relational Database Systems and vice-versa. • Used Spark API over Cloudera Hadoop YARN to perform analytics on data in Hive.• Involved in validating the deployed files using data frames.• Participated in the PI meetings.• Participated in the Sprint Planning meetings.• Followed Agile Software development methodology, with bi-weekly sprints.• Participated in the Release/Demo meetings with the Client