Big Data Engineer
Current• Implemented Spark Scripts using Scala, Spark-SQL to access hive tables into spark for faster processing of data. • Great hands-on experience with Pyspark for using Spark libraries by using python scripting for data analysis.• Work with Avro and Parquet files and convert the data from either format Parsed Semi Structured JSON data and converted to Parquet using Data Frames in PySpark.• Develop Scala scripts, UDFs using both Data frames/SQL and Data sets in Spark for Data Aggregation. • Work on MongoDB database concepts such as locking, transactions, indexes, replication, and schema design.• Develop workflows using Oozie to automate the tasks of loading the data into HDFS & pre-processing with Python.• Use Spark SQL on data frames to access hive tables into spark for faster processing of data.• Develop applications in Spark with Python to transform data.• Migrate Spark applications to Snowpark using Python• Create task, streams, stages, stored-procedures in Snowflake to ingest data• Create Airflow DAGs to monitor and trigger alerts for invalid data• Create ETLs in PySpark to ingest data from heterogenous sources