Senior Data Engineer
Current• Developed applications using spark to implement various aggregation and transformation functions of Spark RDD and Spark SQL. • Worked on DB2 for SQL connection to Spark Scala code to Select, Insert, and Update data into DB. • Used Broadcast Join in SPARK for making smaller datasets to large datasets without shuffling of data across nodes. • Used Oozie Scheduler systems to automate the pipeline workflow and orchestrate the Spark jobs. • Created Spark Streaming jobs using Python to read messages from Kafka & download JSON files from AWS S3 buckets • Used Spark Streaming to receive real-time data from the Kafka and store the stream data to HDFS using Python and NoSQL databases such as HBase and Cassandra• Implemented Spark in EMR for processing Big Data across our One Lake in AWS System • Developed AWS strategy, planning, and configuration of S3, Security groups, IAM, EC2, EMR and Redshift • Developed Spark/Scala, Python for regular expression (regex) project in the Hadoop/Hive environment with Linux/Windows for big data resources. • Data sources are extracted, transformed and loaded to generate CSV data files with Python programming and SQL queries. • Developed data processing applications in Scala using SparkRDD as well as Dataframes using SparkSQL APIs. Environment: Spark, Scala, AWS, Python, Spark SQL, Redshift, PgSQL, Data bricks, Jupiter, Kafka