Data Engineer
Current• Performed data wrangling to clean, transform and reshape the data utilizing panda’s library. Analyzed data using SQL, R, Java, Scala, Python, Apache Spark and presented analytical reports to management and technical teams.• Implemented public segmentation using unsupervised machine learning algorithms by implementing K-means algorithm by using PySpark using data munging.• Experience in Machine learning using NLP text classification, churn prediction using Python.• Lead discussions with users to gather business processes requirements and data requirements to develop a variety of conceptual, logical and Physical Data models.• Expertise in Business intelligence and Data Visualization tools like Tableau.• Handled importing data from various data sources, performed transformations using Hive, MapReduce and loaded data into HDFS.• Used R and Python for Exploratory Data Analysis to compare and identify the effectiveness of the data.• Used Python, R, SQL to create statistical algorithms involving Multivariate Regression, Linear Regression, Logistic