Senior Data Engineer
CurrentDesigned and implemented data ingestion pipelines using Hadoop and Spark to extract, transform, and load data from various sources, including structured, semi-structured, and unstructured data. Implemented data processing and streaming pipelines using Kafka to enable real-time data processing and analytics and designed and implemented custom Kafka connectors to integrate with other systems. Designed and developed custom Power BI and Tableau dashboards to provide real-time insights and visualizations and integrated them with the data processing pipelines using REST APIs and webhooks. Deployed data processing and analytics systems on AWS and Azure using services such as EMR, HDInsight, and Data Factory, and implemented security measures such as VPCs, NSGs, and RBAC. Implemented change data capture (CDC) mechanisms to capture and process real-time data changes, improving data accuracy and timeliness. Developed and maintained data models and architectures using SQL and NoSQL databases such as MySQL, MongoDB, and Cassandra, and optimized query performance using indexing and partitioning techniques.Designed and implemented data warehousing systems using Snowflake, Redshift, and BigQuery, and optimized query performance using techniques such as star schema design and query optimization.Conducted data modelling and schema design using ER diagrams and UML and implemented schema evolution and data partitioning to enable scalability and flexibility.Developed and maintained disaster recovery and business continuity plans for data processing and analytics systems and conducted regular disaster recovery tests and simulations to ensure system availability and readiness.Developed and maintained a data lake architecture to store and process large volumes of data, enabling efficient data storage and processing.