Data Engineer
Current• Built and architected multiple data pipelines and end-to-end ETL processes for data ingestion and transformation in GCP and Migrated on-prem Hadoop systems to Google Cloud Platform.• Deployed Spark and Scala code in a Hadoop cluster running on GCP and developed spark job with partitioned RDD (like hash, range, custom) for faster processing.• Stored data files in Google Cloud S3 Buckets and used GCP services like Dataproc and Big Query for data processing.• Proficient in various Google Cloud components, Container Builder, GCP client libraries, and Cloud SDKs.• Experienced with Azure services, including Data Lake, Data Lake Analytics, SQL Database, Synapse, Databricks, Data Factory, Logic Apps, and SQL Data Warehouse.• Automated the deployment and administration of GKE clusters using Terraform and Kubernetes.• Integrated GKE with GCP’s ecosystem to leverage GCP's networking, storage, and monitoring capabilities, for deploying data engineering solutions in a containerized environment.• Used GKE for high availability and reliability, which makes it a robust choice for deploying and managing data-intensive workloads in a containerized environment.