Senior Data Engineer
Current• Designed and managed scalable storage solutions using GCP Cloud Storage, optimizing data organization, retrieval, and lifecycle management for cost-effective handling of large-scale structured and unstructured data.• Optimized complex queries on BigQuery for large-scale datasets, reducing processing time and enhancing performance.• Integrated machine learning models using Vertex-AI for automated predictions and insights.• Developed scalable data lakes using GCP Cloud Storage for managing both structured and unstructured data.• Orchestrated data workflows with Apache Airflow on Google Composer to automate and streamline ETL processes.• Designed real-time data streaming pipelines using Pub/Sub to facilitate seamless data ingestion and integration.• Built and optimized data pipelines using GCP services such as BigQuery, Cloud Storage, and Composer, improving data processing efficiency.• Integrated predictive models into data workflows using Vertex-AI, enhancing real-time analytics and forecasting accuracy.• Developed real-time data ingestion pipelines with Pub/Sub, enabling seamless integration of event data into analytics systems.reducing operational overhead by 35%.• Optimized data processing workflows using PySpark, enhancing the speed and efficiency of handling large-scale datasets and improving overall pipeline performance.• Built robust ETL pipelines with PySpark, enabling seamless data transformations and ensuring data consistency across distributed systems.• Developed scalable data applications in PySpark to process high-volume data in real-time, achieving faster processing times and greater data accuracy.• Developed data processing applications in Scala for high-performance data transformations, enabling faster, more efficient handling of large datasets.• Implemented complex algorithms using Scala to streamline data processing workflows, improving computational efficiency and reducing execution times.