Lead Data Engineer/Architect
Current•Designed the Big Data Architecture and setup the ETL pipelines using Apache PySpark to manage, store, process data and used Apache Airflow for workflow orchestrations.•Built & deployed the data engineering workflows, machine learning dashboards & analytics, using data bricks and worked on the cloud architectural design. •Led the Enterprise Data Strategy to provide integrated solutions for orchestrating & delivering automations through the Information management team.•Developed prototypes to lead the developers, by providing a working integration of the microservices based architecture.•Collaborated with data scientists, product management & data engineers on Asset management projects to provide guidelines for the effective usage of engineered assets within the portfolio management groups.•Developed the big data architecture & advanced analytics modeling environment for streamlining the data modeling process and speed time-to-decision across the business.•Developed Spark applications using Spark-SQL in Data bricks for data extraction, transformation, and aggregation from multiple file formats for analyzing and transforming the data to gain insights.•Leveraged AWS EMR to transform and move large volumes of data into AWS data stores such as Amazon Simple Storage Service (S3).•Developed Kafka producer and consumers for message handling and executed spark jobs on AWS EMR using data stored in S3 buckets.•Integrated Apache Airflow with AWS to configure and build large scale data pipelines and monitor multi-stage workflows. •Developed snowflake procedures for executing branching and looping and performed data quality issues analysis using Snow SQL by building analytical data warehouses on Snowflake.