Data Engineer
• Assisted in developing and optimizing a scalable data pipeline using PySpark on AWS EMR to handle large datasets efficiently. Contributed to a 20% improvement in data processing speed, enabling quicker access to critical data for business analytics and decision-making.• Collaborated in automating data ingestion processes by setting up workflows in Apache Airflow, reducing manual data handling and increasing data processing efficiency by 20%. This improvement streamlined the ETL process and helped ensure data availability in real time.• Supported the optimization of data models within Snowflake, applying best practices in data warehousing and Snowflake SQL to improve query performance. This optimization led to a 15% increase in analyst productivity by reducing query response times.• Assisted in writing and maintaining SQL queries for data extraction and reporting using SQL Server and PostgreSQL. These queries provided valuable insights to teams by enabling efficient data analysis and reporting.• Automated routine data quality checks using Python, which reduced the time spent on data validation and cleansing by 30%. This ensured accurate and reliable data for reporting and analytics purposes.