Big Data Developer
St Louis, Missouri, United States
• Developed Spark applications using Pyspark and Spark-SQL for data extraction, transformation and aggregation from multiple file formats.• Authoring Python (PySpark) Scripts for custom UDF’s for Row/ Column manipulations, merges, aggregations, stacking, data labeling and for all Cleaning and conforming tasks. Migrate data from on-premises to AWS storage buckets.• Developed frameworks and processes to analyze unstructured information. Assisted in Azure Power BI architecture design.• Created a Python process hosted on Elastic Beanstalk to load the Redshift database daily from several source.• Developed SQL scripts to Upload, Retrieve, Manipulate and handle sensitive data (National Provider Identifier Data I.e. Name, Address, SSN, Phone No) in Teradata, SQL Server Management Studio and Snowflake Databases for the Project.• Experienced in running query using Impala and used BI tools to run ad-hoc queries directly on Hadoop.• Experienced in working with various kinds of data sources such as Teradata and Oracle. Successfully loaded files to HDFS from Teradata, and load loaded from hdfs to hive and impala• Worked on to retrieve the data from FS to S3 using spark commands• Built S3 buckets and managed policies for S3 buckets and used S3 bucket and Glacier for storage and backup on AWS.• Provided access to data necessary to perform analysis on scheduling, pricing, bus bunching and performance. Queries in the Redshift environment performed 100-1000x faster than in legacy environments.• Transform and analyze the data using Pyspark, HIVE, based on ETL mappings• Developed Unix Shell Scripts and automated them using CRON job scheduler.• Experienced with machine learning algorithm such as logistic regression, random forest, XGboost, KNN, SVM, neural network, linear regression, lasso regression and k - means• Implemented Statistical model and Deep Learning Model (Logistic Regression, XGboost, Random Forest, SVM, RNN, and CNN).