Data Engineer
CurrentGenerating a Big Data solution using Spark and Scala to analyze the MovieLens dataset (1 million records), answering analytical questions on movie ratings, genres, and user behavior.Leveraging Spark RDD, Spark SQL, and Spark DataFrame to process and analyze large datasets, improving data processing speed; utilizing Cassandra for efficient data storage.Exposing the solution as an API, implementing logging for tracking over 100+ actions, and managing code on GitHub, ensuring modularity and following best practices.Deploying the solution on cloud platforms such as AWS or GCP, optimizing performance and latency using Ops Pipeline tools like Jenkins and Azure DevOps.