Senior Data Engineer
Current Used Sqoop to import data into HDFS/Hive from multiple relational databases, performed operations, and exported the results back. Got involved in migrating the on-prem Hadoop system to using GCP (Google Cloud Platform).Demonstrated expertise in leveraging GCP services such as Compute Engine, Kubernetes Engine, Cloud Storage, BigQuery, and Cloud SQL for seamless migration and efficient operation of workloads. Extensively used Spark Streaming to perform the analysis of credit department data on real-time regular window time intervals coming from sources like Kafka. Performed Spark join optimizations, troubleshooting, monitored, and wrote efficient codes using Scala. Used big data tools Spark (Pyspark, SparkSQL) to conduct real-time analysis for banking transactions. Performed Spark transformations and actions on large datasets. Implemented Spark SQL to perform complex data manipulations, and to work with large amounts of structured and semi-structured data stored in a cluster using Data Frames/Datasets. Working with big data tools: Hadoop, Spark, Kafka, etc. Support existing GCP Data Management implementations. Created GCP Big Query authorized views for row-level security or exposing the data to other teams. Exploring with Spark to improve the performance and optimization of the existing algorithms in Hadoop using Spark SQL, Data Frame, pair RDDs, and Spark YARN. Experience with terraform scripts which automate the step execution in EMR to load the data to Scylla DB. De-normalizing the data as part of transformation which is coming from Netezza and loading it to No Sql Databases and MySQL. Loading data into snowflake tables from the internal stage using SnowSQL. Used import and export from the internal stage (snowflake). Implemented metadata and documentation management practices for Erwin modeling and data cataloging, ensuring accurate and up-to-date documentation of the data environment