Big Data Engineer
Current• Data extraction from APIs and Databases, stored in Lake-house, and processed using PySpark and SparkSQL. Managed data through Raw, Landing, Curated, and Serving Layers. • Utilised Fivetran as an ELT tool to establish seamless connectivity between data sources, enabling efficient data extraction and loading. • Interpreted and migrated old legacy stored procedures into optimised PySpark code, incorporating dynamic functionality to enhance scalability and performance.• Data transformation using Notebooks or Data Build Tool (DBT) and loading it into a Fabric Lake-house or Snowflake warehouse. • Implemented schema validation and data quality checks with Great Expectations. Managed Slowly Changing Dimensions (SCDs) for historical data. • Developed control tables and logging mechanisms for pipeline status and quality checks. Status tables creation for performance monitoring. • Designed and implemented data models using Draw.io and Oracle Developer Data Modeler, facilitating efficient database schema design and management. • Process and coding documentation in Markdown / pydoc, maintained in DevOps for version control and collaboration. • Designed and deployed interactive dashboards and reports using Power BI for data insights and visualisation.• Conducted POCs to modernise data pipelines on Azure Fabric, contributed to designing a star-schema-based data warehouse and data lake, and developed transformation workflows with Spark Notebooks. Collaborated on creating semantic layers, ensuring data quality with Great Expectations, and integrating solutions with existing systems.