Data Engineer Senior Consultant/ Tech Lead
Current- Implemented ETL workflows to handle real-time and scheduled data ingestion, transformation, and loading, ensuring seamless integration with various data sources and destinations- minimizing manual efforts- Engineered scalable data pipelines using PySpark’s DataFrame API and SQL capabilities to perform complex transformations and aggregations on large datasets- Auto-scaled and optimized clusters to enhance pipeline performance and reliability by configuring auto-scaling policies and selecting appropriate instance types based on workload requirements. Leveraged Databricks' cluster pools and spot instances to efficiently manage resources and reduce costs- Optimized data performance through StreamSets configuration and leveraging Snowflake’s parallel processing capabilities, resulting in faster data processing time and improved system efficiency- Developed ETL processes using Apache Spark on Databricks to ingest data from various sources into data lakes and warehouses and load into AWS S3 buckets for further analysis- Facilitated sprint meetings and progress reviews for a team of 4 engineers using JIRA, overseeing task creation and assignment, addressing roadblocks, and ensuring alignment with project objectives. Managed timelines and made necessary adjustments to keep the project on track, while fostering team collaboration and communication- Coordinated with stakeholders to gather requirements, develop project plans, and track progress, ensuring alignment with business goals- Engaged with end-users to gather feedback on data solutions and incorporated user insights into the design of ETL processes and dashboards, resulting in improved user satisfaction and adoption