Hadoop Developer
Current• Evaluated business requirements and prepared detailed specifications that follow project guidelines required to develop written programs.• Exploring with the Spark improving the performance and optimization of the existing algorithms in Hadoop using Spark Context, Spark-SQL, Data Frame, Pair RDD's, Spark YARN.• Managed and reviewed Hadoop log files to identify issues when job fails and used HUE for UI based pig script execution, Oozie scheduling.• Involved in creating data-lake by extracting customer's data from various data sources to HDFS which include data from Excel, databases, and log data from servers.• Developed Python code to gather the data from HBase and designs the solution to implement using PySpark.• Developed PySpark code to mimic the transformations performed in the on-premise environment and analyzed the SQL scripts and designed solutions to implement using PySpark.• Automated workflows using shell scripts pull data from various databases into Hadoop and developed scripts to automate the process and generate reports.• Created detailed AWS Security groups which behaved as virtual firewalls that controlled the traffic allowed reaching one or more AWS EC2 instances.• Designed multiple Python packages that were used within a large ETL process used to load 2TB of data from an existing Oracle database into a new PostgreSQL cluster.• Deploy and configured cloud AWS EC2 for client websites moving from self-hosted services for scalability purposes and work with multiple teams to provision AWS infrastructure for development and production environments.• Designed number of partitions and replication factor for Kafka topics based on business requirements and worked on migrating MapReduce programs into Spark transformations using Spark and Scala, initially done using python (PySpark).System (HDFS) on Amazon EMR cluster by setting up the Spark Core for analysis work.