Data Engineer
Current● Designing a robust data warehouse using dimensional modeling, seamlessly integrating data from Google Ads API and Mailchimp, encompassing all campaign operations to facilitate informed decision-making● Employing Dbt for precise data transformation tailored to client needs, while harnessing Airbyte to seamlessly ingest data from Google Ads, Mailchimp, and existing client data warehouse, optimizing business insights and operations ● Developed a robust ETL pipeline with Python and AWS Lambda to automate data transfer of watch csv data from S3 to a PostgreSQL RDS database, optimizing data handling efficiency and accuracy.● Leveraged PySpark using Databricks, EMR and AWS Glue to efficiently transform unstructured, semi-structured (JSON ,CSV), and structured data from API, relational databases(RDS, PostgreSQL), data lake (S3), optimizing ETL pipelines within Airflow for comprehensive data integration● Efficiently designed a Python-based ETL pipeline on EC2, integrating Google Drive API to seamlessly transfer and store 3 GB of company data in Amazon S3, saving approximately 20 hours of manual uploads per week● Designed PostgreSQL tables and schemas through data modeling using SQL, facilitating seamless data access and retrieval for data scientists and developers, enhancing data usability and analytics capabilities● Successfully designed and implemented logging strategies in ETL data pipeline in AWS by leveraging AWS CloudWatch to monitor and track processes, ensuring real-time visibility into application performance● Implemented effective version control practices using Git to streamline code review processes, ensuring efficient and error-free integration of changes into the codebase● Collaborated closely with data scientists, data analyst, project managers, and development teams within agile sprints, actively participating in data-related tasks and presenting outcomes on a weekly basis