Project Lead
Current• Responsible for coding Map Reduce program, Hive queries, testing and debugging the Map Reduce programs.• Responsible for Installing, Configuring and Managing of Hadoop Cluster spanning multiple racks.• Developed Pig Latin scripts in the areas where extensive coding needs to be reduced to analyze large data sets.• Used Sqoop tool to extract data from a relational database into Hadoop.• Involved in performance enhancements of the code and optimization by writing custom comparators and combiner logic.• Worked closely with data warehouse architect and business intelligence analyst to develop solutions.• Good understanding of job schedulers like Fair Scheduler which assigns resources to jobs such that all jobs get, on average, an equal share of resources over time and an idea about Capacity Scheduler.• Responsible for performing peer code reviews, troubleshooting issues and maintaining status report.• Involved in creating Hive Tables, loading with data and writing Hive queries, which will invoke and run MapReduce jobs in the backend.• Involved in loading data from UNIX file system to HDFS