Senior Data Engineer
CurrentDesigned and set up Enterprise Data Lake on AWS to support storing, processing, analytics, and reporting of large and dynamic datasets using services like S3, EC2, ECS, AWS Glue, SNS, SQS, DMS, and Kinesis.Choose suitable ETL frameworks like Apache Spark or AWS Glue based on project needs and data volume.Configured and maintained PostgreSQL databases for reliable data storage.Developed and managed scalable ETL pipelines on Databricks to process large datasets… Show more Designed and set up Enterprise Data Lake on AWS to support storing, processing, analytics, and reporting of large and dynamic datasets using services like S3, EC2, ECS, AWS Glue, SNS, SQS, DMS, and Kinesis.Choose suitable ETL frameworks like Apache Spark or AWS Glue based on project needs and data volume.Configured and maintained PostgreSQL databases for reliable data storage.Developed and managed scalable ETL pipelines on Databricks to process large datasets efficiently.Developed and managed data pipelines using Databricks on AWS to handle large-scale data processing and analytics.Developed and maintained REST APIs to facilitate seamless communication between microservices and external systems. Developed data processing pipelines using PySpark to handle large-scale datasets efficiently.Designed and implemented Kafka clusters to handle large-scale, real-time data streaming and processing.Developed Python scripts for data transformation and integration, ensuring high performance and reliability across various data sources.Used AWS Databricks to build and manage scalable data pipelines and analytics workflows in the cloud.Used ELT to handle large volumes of data by leveraging the target system's processing power for transformations.Leveraged Kafka API to build scalable event-driven architectures, supporting high-throughput data pipelines.Managed infrastructure as code with Terraform, enabling version control and collaboration on infrastructure changes.Developed and managed workflows using Airflow to automate and schedule data pipelines.Developed interactive dashboards in Power BI to visualize complex data trends and insights, improving stakeholder engagement and data accessibility.Optimized REST API performance by implementing caching, load balancing, and efficient data serialization techniques storage. Integrated PySpark with various data sources, including HDFS, S3, and Hive, to streamline data ingestion and transformation. Show less