Data Engineering Associate Manager
CurrentManaging and leading a team of 26 from freshers to experienced engineers to conceptualize, design and implement lambda architecture to create a ground up data platform with scope of cloud migrations in future. It is designed to handle 800+ millions of records. The biggest achievements in this project were to build generic reusable solution to handle 600+ heterogenous sources and implement SCD 1 and SCD 2 which saved over 5000 hours as compared to conventional designs. Other achievements include moving the team from SLDC practice to Agile practice and completely removing the dependencies on senior resources in the project by introducing specific practices and processes. The solution was built on IBM DataStage, Kafka, NiFi. Databases involved were Oracle, Hive, Teradata & Mongo. Next challenge for this team is to build solutions on Apache Flink, which is a new tech to the team, and I am creating learning path strategy and resources for hands-on practice. Building Solutions to identify Teradata connectors and replace it will BigQuery connectors automatically at enterprise level DataStage as part of Teradata to BigQuery migration. Building configurable and resilient data quality framework to flag records with data issues and filtered them out not to be processed further. Current processing rate of the framework is 110 million records per minute and the solution is built on PL/SQL. Delivered POCs as per client request to build a scalable PubSub data pipeline using python, Google PubSub and BigQuery. Others included were handing SCD data built on Snowflake using Streams and Tasks. Single handedly worked on multiple shell scripts to bring in automation to multiple processes like exporting DataStage information to be fed to data lineage tool and automating deployment from exporting the jobs from dev to deploying in production.Other engagements include building cloud agnostic solutions and create a multi-cloud strategy for stressed exit.