Data Engineer
Current• Involved in importing and exporting data between Hadoop Data Lake and Relational Systems like Oracle, MySQL using Sqoop.• Involved in developing spark applications to perform ELT kind of operations on the data.• Modified existing MapReduce jobs to Spark transformations and actions by utilizing Spark RDDs, Data frames, and Spark SQL API’s• Create and maintain optimal data pipeline architecture in cloud Microsoft Azure using Data Factory and Azure Databricks• Create an Architectural solutions that leverages the best Azure analytics tools to solve our specific need in Chevron use case• Design and present technical solutions to end users in a way the is easy to understand and buy into• Educating client/business users on the pros and cons of various Azure PaaS and SaaS solutions ensuring the most cost-effective approaches are taken into consideration.• Written Templates for Azure Infrastructure as Code using Terraform to build staging and production environments.• Utilized Hive partitioning, Bucketing and performed various kinds of joins on Hive tables• Involved in creating Hive external tables to perform ETL on data that is produced on daily basis.• Design and developed services to persist and rea data from Hadoop, HDFC, hive and writing Java based MapReduce batch jobs using Hortonworks Hadoop data platform.• Designed and deployed data pipelines using Data Lake, Databricks, and Apache Airflow.• Install and configure Apache Airflow for S3 bucket and snowflake data warehouse and created dags to run the Airflow.• Developed Spark applications using PySpark and Spark-SQL for data extraction, transformation, and aggregation from multiple file formats for analyzing & transforming the data to uncover insights into the customer usage patterns.• User Profile and other unstructured data storage using Java and MongoDB