Senior Data Engineer
Current● Created ETL pipelines using Python, Spark, PySpark and Hive to process data and load it into Oracle and S3.● Work on database upgrade activities and create automated jobs to validate data / structures between old and new databases.● Test various processes and support different models and releases for Actimize and load the input data for validations into prod-beta environments to test compatibility and performance.● Provide performance tuning for the existing jobs and redesign if required.● Support upgrade activities by coordinating with multiple teams for setup of environment and services.● Build data pipelines using Data Proc and Dataflow in GCP for ETL related jobs.● Worked on GCP data-proc cloud platform to process huge volume datasets.● Used cloud shell SDK in GCP to configure the services Data Proc, Cloud Storage, BigQuery.● Worked with google data catalog and other google cloud APIs for monitoring, query, and billing related analysis for BigQuery usage.● Carried out data transformation and cleansing using SQL queries, Python and Pyspark.● Coordinating with team and Developed framework to generate Daily adhoc reports and Extracts from enterprise data from BigQuery.● Wrote scripts in Hive SQL for creating complex tables with high performance metrics like partitioning, clustering, and skewing.● Worked on downloading Big Query data into pandas or Spark data frames for advanced ETL capabilities.