Big Data Engineer
CurrentResponsible for implementation, administration and management of Hadoop infrastructuresEvaluation of Hadoop infrastructure requirements and design/deploy solutions (high availability, big data clusters and involved in cluster monitoring and troubleshooting Hadoop issuesWorked with application teams to install OSs and Hadoop updates, patches, version upgrades as requiredHelped maintain and troubleshoot UNIX and Linux environmentAnalyzed and evaluated system security threats and safeguardsDeveloped Pig program for loading and filtering the streaming data into HDFS using Flume.Experienced in handling data from different datasets, join them and preprocess using Pig join operations.Developed HBase data model on top of HDFS data to perform real time analytics using Java API.Developed different kinds of custom filters and handled predefined filters on HBase data using API.Imported and exported data from Teradata to HDFS and vice-versa.Strong understanding of Hadoop eco system such as HDFS, MapReduce, HBase, Zookeeper, Pig, Hadoop streaming, Sqoop, Oozie and HiveImplemented Secondary sorting to sort reducer output globally in map reduce.Implemented data pipeline by chaining multiple mappers by using Chained Mapper.Created Hive Dynamic partitions to load time series dataHandled different types of joins in Hive like Map joins, bucker map joins,sorted bucket map joins.Created tables, partitions, buckets and performed analytics using Hive ad-hoc queries.Experienced import/export data into HDFS/Hive from relational database and Tera data using Sqoop.Handled continuous streaming data from different sources using Flume and set destination as HDFS.Integrated spring schedulers with Oozie client as beans to handle corn jobs.Experience with CDH distribution and Cloudera Manager to manage and monitor Hadoop clusters