Big Data Developer/Hadoop Developer
Current• Programmed in Hive, Spark SQL and Python to streamline the incoming data and build the data pipelines to get the useful insights, and orchestrated pipelines using Azure Data Factory.• Used ORC, Parquet file formats on HDInsight, Azure Blobs and Azure tables to store for raw data.• Implemented Hive tables and HQL Queries for the reports. Written and used complex data type in Hive. Storing and retrieved data using HQL in Hive. Developed Hive queries to analyze reducer output data.• Designed the next generation data architecture for the unstructured data.Environment: CDH4, HDFS, Map Reduce, Hive, Oozie, Java, Amazon EMR, Amazon S3, Amazon Glue, CloudWatch, Amazon Kinesis,, PIG, Shell Scripting, Maven, HDP 2.6, Linux, HUE, Airflow, Sqoop, Akka, Tableau , Flume, Snowflake, NoSQL, DB2, and Oracle.