Data Engineer
CurrentData EngineerAs a Data Engineer, I specialize in designing and maintaining scalable, efficient, and reliable data pipelines that empower data-driven decision-making. My day-to-day responsibilities include:Data Quality and Integrity: Conducting comprehensive data sanity checks using SQL and Looker Studio to ensure the accuracy and reliability of datasets.Pipeline Development: Building robust and automated data pipelines using Apache Airflow and various operators, integrating complex workflows seamlessly.Programming Expertise: Leveraging Scala, Python, PySpark, and Spark to create scalable and efficient pipelines for diverse data processing needs.Cloud Solutions: Developing and deploying cloud-native solutions on Google Cloud Platform (GCP), working extensively with Dataproc, Google Cloud Storage (GCS), BigQuery, and Firestore.Data Warehousing: Managing and optimizing data warehousing solutions using ClickHouse to enable high-performance analytical queries.Collaboration: Partnering with cross-functional teams, including data analysts, scientists, and business stakeholders, to deliver actionable insights and support organizational goals.Optimization and Monitoring: Continuously improving pipeline performance, ensuring fault-tolerance, and monitoring data workflows for efficiency and reliability.Innovation and Best Practices: Staying updated with the latest industry trends in data engineering, implementing best practices, and exploring emerging technologies to drive innovation in data workflows.