Senior Data Engineer
CurrentOver 10 years of experience as GCP Data Engineer with demonstrated expertise in building and deploying data pipelines using open-source Hadoop based technologies such as Apache Spark, Hive, Hadoop, Python and PySpark.Hands on Experience in developing Spark applications using PySpark Data Frame, RDD and Spark SQL.Working with GCP cloud using in GCP Cloud storage, DataProc, Data Flow, Big Query, Cloud Composer and Cloud Pub/Sub.Expert in working with cloud PUB/SUB to replicate data real-time from source system to GCP Big Query.Good knowledge on GCP service accounts, billing projects, authorized views, datasets, GCS buckets and gsutil commands.Experienced in building and deploying Spark applications on Hortonworks Data Platform and AWS EMR.Experienced in working with AWS services such as – EMR, S3, EC2, IAM, Lambda, Cloud Formation, Cloud Watch.Worked on structured and semi structured data storage formats such as Parquet, ORC, CSV, JSON.Hands on experience on Google Cloud Platform (GCP in all the big data products Big Query, Cloud Data Proc, Google Cloud Storage, Composer (AirFlow as a service)Hands on experience working in GCP services like Big Query, Cloud Storage (GCS), cloud function, cloud dataflow, Pub/sub, Cloud Shell, GSUTIL, Big Query, Data Proc, Operations Suite (Stack driver).