Sr. Data Enginner
Current• Managed end-to-end Big Data flow within the application, including data ingestion from upstream sources to HDFS, as well as processing and analysis of data in HDFS.Worked extensively with PySpark on EMR and Databricks, leveraging the Hadoop ecosystem for efficient data processing.Utilized Spark core, Spark SQL, and Spark Streaming for real-time data processing and analysis.Designed and implemented a learner data model that receives real-time data from Kafka and persists it in Cassandra, enabling real-time data storage and analysis.Utilized various AWS cloud components, such as EMR, S3, EC2, Athena, RDS, Dynamo DB, Lambda, IAM and Glue