Senior Data Engineer
•Proficient in crafting and designing multiple data pipelines, overseeing the complete ETL and ELT process for data ingestion and transformation within the Google Cloud Platform (GCP).•Successfully set up a Continuous Delivery pipeline using Docker and GitHub, streamlining the deployment process.•Re-platformed AWS EMR-based Spark and Hadoop jobs to GCP Dataproc, leveraging GCP's autoscaling and managed cluster features for optimized resource management and cost control.•Transitioned logging and monitoring solutions from AWS CloudWatch to Google Cloud Monitoring and Logging to maintain visibility into the migrated data pipelines and ensure performance metrics are equal.•Developed, deployed, and managed results utilizing Spark and Scala code within a Hadoop cluster hosted on GCP.•Skilled in leveraging Google Cloud components, Google Container Builders, GCP client libraries, and Cloud SDKs to architect and execute data solutions.•Hands-on experience with Google Cloud Function, using Python to load data into BigQuery from incoming CSV files in GCS buckets. Also, proficient in processing and loading both bounded and unbounded data from Google Pub/Sub topics to BigQuery via Cloud Dataflow.•Utilized Spark and Scala APIs to assess the performance of Spark in comparison to Hive and SQL. A•Proficiently stored data in the GCP BigQuery Target Data Warehouse, making it available for various business teams according to their specific use cases.•Successfully deployed applications to GCP using Spinnaker, leveraging rpm-based packages.•Architected several Directed Acyclic Graphs (DAGs) to automate ETL pipelines for seamless data processing.• Leveraged Amazon EMR clusters for processing large-scale datasets, optimizing for performance and cost-efficiency before transitioning to GCP.• Integrated AWS services with GCP components, such as Cloud Dataflow and BigQuery, for a hybrid cloud solution during migration.